1.8.9b10 — a free model after credits run out
An opt-in preview on the 1.8.9 line. Stable remains 1.8.8.
Managed text model
New Agents and llm_do() calls use co/llama by default. It runs Llama 3.1
8B on ConnectOnion's GPU through oo-api v0.1.22 and has a zero-dollar token
price. co/gemma is an explicit free alternative. Paid models remain available
when selected by name. Audio transcription keeps its Gemini default because
this Llama route accepts text.
When a paid call reaches a credit limit, oo-api includes free_model in the
error and the SDK names that route in its message. Calls to either free model
can continue with an exhausted balance.
The deployed API returned co/llama, a valid answer, and usage.cost_usd: 0.0
for a fresh account that omitted model. The account balance stayed at $5. A
second production call returned a get_weather(city=Sydney) tool call, and the
SDK's default llm_do() returned an answer through the same API.
Install
python -m pip install --upgrade 'connectonion==1.8.9b10'
Known limits
- The shared GPU proxy admits one request at a time. Busy requests receive HTTP 429 and may need a retry.
- Ollama currently uses a 4,096-token context and caps output at 1,024 tokens.
- This preview is opt-in; the default stable package remains 1.8.8.
ConnectOnion