# 1.8.9b10 — a free model after credits run out

An opt-in preview on the 1.8.9 line. Stable remains 1.8.8.

## Managed text model

New Agents and `llm_do()` calls use `co/llama` by default. It runs Llama 3.1
8B on ConnectOnion's GPU through oo-api v0.1.22 and has a zero-dollar token
price. `co/gemma` is an explicit free alternative. Paid models remain available
when selected by name. Audio transcription keeps its Gemini default because
this Llama route accepts text.

When a paid call reaches a credit limit, oo-api includes `free_model` in the
error and the SDK names that route in its message. Calls to either free model
can continue with an exhausted balance.

The deployed API returned `co/llama`, a valid answer, and `usage.cost_usd: 0.0`
for a fresh account that omitted `model`. The account balance stayed at $5. A
second production call returned a `get_weather(city=Sydney)` tool call, and the
SDK's default `llm_do()` returned an answer through the same API.

## Install

```sh
python -m pip install --upgrade 'connectonion==1.8.9b10'
```

## Known limits

- The shared GPU proxy admits one request at a time. Busy requests receive
  HTTP 429 and may need a retry.
- Ollama currently uses a 4,096-token context and caps output at 1,024 tokens.
- This preview is opt-in; the default stable package remains 1.8.8.
