ConnectOnionConnectOnion

1.8.9b10 — a free model after credits run out

An opt-in preview on the 1.8.9 line. Stable remains 1.8.8.

Managed text model

New Agents and llm_do() calls use co/llama by default. It runs Llama 3.1 8B on ConnectOnion's GPU through oo-api v0.1.22 and has a zero-dollar token price. co/gemma is an explicit free alternative. Paid models remain available when selected by name. Audio transcription keeps its Gemini default because this Llama route accepts text.

When a paid call reaches a credit limit, oo-api includes free_model in the error and the SDK names that route in its message. Calls to either free model can continue with an exhausted balance.

The deployed API returned co/llama, a valid answer, and usage.cost_usd: 0.0 for a fresh account that omitted model. The account balance stayed at $5. A second production call returned a get_weather(city=Sydney) tool call, and the SDK's default llm_do() returned an answer through the same API.

Install

python -m pip install --upgrade 'connectonion==1.8.9b10'

Known limits

  • The shared GPU proxy admits one request at a time. Busy requests receive HTTP 429 and may need a retry.
  • Ollama currently uses a 4,096-token context and caps output at 1,024 tokens.
  • This preview is opt-in; the default stable package remains 1.8.8.

View Markdown source

Star us on GitHub

If ConnectOnion saves you time, a ⭐ goes a long way — and earns you a coffee chat with our founder.