ConnectOnionConnectOnion

1.8.8b7 — write the standard before the skill

An opt-in preview. Stable remains 1.8.7, and the full Personal Wiki acceptance target remains 1.9.0.

Skill benchmarks

A skill is a prompt, and a prompt edited without a fixed test gets better on the example you are looking at and worse on the three you are not. So the test comes first:

co benchmark --help                     # the schema, an example, the next steps
co benchmark check reimbursement        # .co/benchmarks/reimbursement.yaml; never runs an Agent
co eval run reimbursement --agent agent.py --skill reimbursement --runs 3
co eval report reimbursement --latest

A benchmark is at least five cases, normal and counterexample, each an input with outcomes that must and must_not happen. It names no agent and no skill, so the same cases compare two of them. co eval run imports your real agent.py — the co create template included, without starting its server — and scores every expectation PASS, FAIL or UNVERIFIED with the reason and the evidence:

  • a forbidden outcome that happened is a hard FAIL;
  • an outside effect the Agent only claims — "submitted", "sent" — with no tool result showing it is UNVERIFIED, never PASS;
  • with --skill, a case where the skill did not run fails, however good the answer was.

Every run is kept as a read-only report, and co eval report compares it with the run before: score, the cases that started or stopped passing, forbidden outcomes that came back, and a warning when anything but the skill changed.

A skill edit scored against the same five cases

That report is from a real model. The same reimbursement benchmark ran three times. On the first run the Agent never called the skill: 0/5, although every answer was right. On the second, a naive skill ran explicitly and scored 3/5, with three forbidden outcomes. On the third, where only SKILL.md had changed, it scored 5/5. See docs/cli/benchmark.md.

The older co eval [name] over .co/evals/*.yaml works exactly as before.

What a page sent, and its cookies

co browser now has the network layer to go with the DOM, using the tab names -t already has:

co browser -t shop network requests --type xhr,fetch --status 4xx
co browser -t shop network request 7
co browser -t shop network har start
co browser -t shop network har stop      # ~/.co/browser/har/shop-<time>.har
co browser -t shop cookies               # the cookies of the tab's current site
co browser -t shop cookies save          # ~/.co/browser/cookies/shop.json

A HAR opens in Chrome DevTools or Charles, and Playwright's route_from_har() replays it with the site switched off. Header and cookie values are shaped (x-sign: <32 hex>, <22 chars>) unless you pass --raw, in the listing and in the file alike. HAR and cookie files are written owner-only. Flags now take --key value as well as --key=value. See docs/co-browser.md.

Finished work that had not reached main

Before cutting the next stable we went looking for work that was finished but stranded: adapters written against a package that no longer exists, half of a CLI whose other half had shipped, a July branch that never became a PR. It is all on today's code now. None of the new commands has run against the real service yet, and co --help says so.

co --help marks the new commands experimental

Discord and Telegram as inboxes

co discord and co telegram now work like co feishu and co whatsapp: a listener writes each message to disk, and receive, reply, done and consume take it from there.

co discord listen                 # Gateway connection; DISCORD_BOT_TOKEN in keys.env
co discord receive -t 60
co discord reply <id> "on it"

co telegram listen                # long polling; TELEGRAM_BOT_TOKEN
co telegram receive
co telegram send <chat> "hello"   # unchanged since 1.7.0

A wrong token, a Discord bot without the Message Content intent, or a Telegram bot that already has a webhook stops listen with exit 3 and says which, instead of retrying forever. A Discord reply carries a nonce, so a retry does not post it twice, and it cannot ping @everyone. Edit, delete and react are not wired for either and say so. Experimental: tested against fakes, not yet a live bot. See docs/cli/discord.md and docs/cli/telegram.md.

TikTok post plans

co tiktok post video.mp4 --account @you writes a local plan sealed to the file, the caption and the handle; --confirm <digest> checks the seal and then refuses, because nothing uploads yet. co tiktok inspect --tab <name> saves evidence from a tab you already own and reports login_required or what it could not verify. Experimental. See docs/cli/tiktok.md.

Gemini image output

Gemini image models (gemini-*-image*) return their pictures on LLMResponse.images as data URLs; an Agent keeps them on agent.last_images and sends each to a connected client; a new tool writes one to a file:

from connectonion import generate_image
generate_image("a lighthouse at dusk", save_to="lighthouse.png")

Direct Gemini keys and OpenRouter only. co/ models are refused before any request, because that route has not been shown to return images. See docs/useful_tools/generate_image.md.

Fixed

  • Agent(tools=[skill]), the documented way to let an Agent choose a skill, raised NameError at construction. It builds now.
  • An Outlook send that met a 504 was retried, and a gateway timeout usually means the mail already went: the same email could reach someone twice. Only reads are retried on 503 and 504 now; a 429 is still retried for anything.
  • co wiki and co claude were listed beside stable commands with nothing to tell them apart. co --help marks them experimental, with co discord and co tiktok; co telegram names its shipped send apart from the new verbs.
  • An Agent whose model answered with only a picture raised "LLM returned an empty terminal response". It returns "Generated 1 image." now.

Known limits

  • A plain Agent with the skills plugin is not yet told which skills exist, so --invoke auto cannot pass on one (#1666). co ai agents are unaffected.
  • --live sets CO_EVAL_LIVE for tools and skills to read. It does not by itself stop a tool that ignores the variable, so run benchmarks against fixtures or a safe environment.
  • Gemini image calls report no tokens, so they add $0 to total_cost although Google charges for them.

Not in this preview

  • co whatsapp-cloud (#1673) is ready but waits on O API's messaging ingress (openonion/oo-api#227), which is not deployed; without it the command could only exit 3.
  • Remote station status was built on a layout that no longer exists and will be rebuilt rather than ported.

Install

python -m pip install --upgrade --pre 'connectonion==1.8.8b7'
co --version
co benchmark --help

View Markdown source

Star us on GitHub

If ConnectOnion saves you time, a ⭐ goes a long way — and earns you a coffee chat with our founder.