1.8.8b7 — write the standard before the skill
An opt-in preview. Stable remains 1.8.7, and the full Personal Wiki acceptance target remains 1.9.0.
Skill benchmarks
A skill is a prompt, and a prompt edited without a fixed test gets better on the example you are looking at and worse on the three you are not. So the test comes first:
co benchmark --help # the schema, an example, the next steps
co benchmark check reimbursement # .co/benchmarks/reimbursement.yaml; never runs an Agent
co eval run reimbursement --agent agent.py --skill reimbursement --runs 3
co eval report reimbursement --latest
A benchmark is at least five cases, normal and counterexample, each an input
with outcomes that must and must_not happen. It names no agent and no
skill, so the same cases compare two of them. co eval run imports your real
agent.py — the co create template included, without starting its server —
and scores every expectation PASS, FAIL or UNVERIFIED with the reason and the
evidence:
- a forbidden outcome that happened is a hard FAIL;
- an outside effect the Agent only claims — "submitted", "sent" — with no tool result showing it is UNVERIFIED, never PASS;
- with
--skill, a case where the skill did not run fails, however good the answer was.
Every run is kept as a read-only report, and co eval report compares it with
the run before: score, the cases that started or stopped passing, forbidden
outcomes that came back, and a warning when anything but the skill changed.

That report is from a real model. The same reimbursement benchmark ran three
times. On the first run the Agent never called the skill: 0/5, although every
answer was right. On the second, a naive skill ran explicitly and scored 3/5,
with three forbidden outcomes. On the third, where only SKILL.md had changed,
it scored 5/5. See docs/cli/benchmark.md.
The older co eval [name] over .co/evals/*.yaml works exactly as before.
What a page sent, and its cookies
co browser now has the network layer to go with the DOM, using the tab
names -t already has:
co browser -t shop network requests --type xhr,fetch --status 4xx
co browser -t shop network request 7
co browser -t shop network har start
co browser -t shop network har stop # ~/.co/browser/har/shop-<time>.har
co browser -t shop cookies # the cookies of the tab's current site
co browser -t shop cookies save # ~/.co/browser/cookies/shop.json
A HAR opens in Chrome DevTools or Charles, and Playwright's
route_from_har() replays it with the site switched off. Header and cookie
values are shaped (x-sign: <32 hex>, <22 chars>) unless you pass --raw,
in the listing and in the file alike. HAR and cookie files are written
owner-only. Flags now take --key value as well as --key=value. See
docs/co-browser.md.
Finished work that had not reached main
Before cutting the next stable we went looking for work that was finished but
stranded: adapters written against a package that no longer exists, half of a
CLI whose other half had shipped, a July branch that never became a PR. It is
all on today's code now. None of the new commands has run against the real
service yet, and co --help says so.

Discord and Telegram as inboxes
co discord and co telegram now work like co feishu and co whatsapp:
a listener writes each message to disk, and receive, reply, done and
consume take it from there.
co discord listen # Gateway connection; DISCORD_BOT_TOKEN in keys.env
co discord receive -t 60
co discord reply <id> "on it"
co telegram listen # long polling; TELEGRAM_BOT_TOKEN
co telegram receive
co telegram send <chat> "hello" # unchanged since 1.7.0
A wrong token, a Discord bot without the Message Content intent, or a Telegram
bot that already has a webhook stops listen with exit 3 and says which,
instead of retrying forever. A Discord reply carries a nonce, so a retry does
not post it twice, and it cannot ping @everyone. Edit, delete and react are
not wired for either and say so. Experimental: tested against fakes, not yet a
live bot. See docs/cli/discord.md and
docs/cli/telegram.md.
TikTok post plans
co tiktok post video.mp4 --account @you writes a local plan sealed to the
file, the caption and the handle; --confirm <digest> checks the seal and then
refuses, because nothing uploads yet. co tiktok inspect --tab <name> saves
evidence from a tab you already own and reports login_required or what it
could not verify. Experimental. See docs/cli/tiktok.md.
Gemini image output
Gemini image models (gemini-*-image*) return their pictures on
LLMResponse.images as data URLs; an Agent keeps them on agent.last_images
and sends each to a connected client; a new tool writes one to a file:
from connectonion import generate_image
generate_image("a lighthouse at dusk", save_to="lighthouse.png")
Direct Gemini keys and OpenRouter only. co/ models are refused before any
request, because that route has not been shown to return images. See
docs/useful_tools/generate_image.md.
Fixed
Agent(tools=[skill]), the documented way to let an Agent choose a skill, raisedNameErrorat construction. It builds now.- An Outlook send that met a 504 was retried, and a gateway timeout usually means the mail already went: the same email could reach someone twice. Only reads are retried on 503 and 504 now; a 429 is still retried for anything.
co wikiandco claudewere listed beside stable commands with nothing to tell them apart.co --helpmarks them experimental, withco discordandco tiktok;co telegramnames its shippedsendapart from the new verbs.- An Agent whose model answered with only a picture raised "LLM returned an empty terminal response". It returns "Generated 1 image." now.
Known limits
- A plain
Agentwith the skills plugin is not yet told which skills exist, so--invoke autocannot pass on one (#1666).co aiagents are unaffected. --livesetsCO_EVAL_LIVEfor tools and skills to read. It does not by itself stop a tool that ignores the variable, so run benchmarks against fixtures or a safe environment.- Gemini image calls report no tokens, so they add $0 to
total_costalthough Google charges for them.
Not in this preview
co whatsapp-cloud(#1673) is ready but waits on O API's messaging ingress (openonion/oo-api#227), which is not deployed; without it the command could only exit 3.- Remote station status was built on a layout that no longer exists and will be rebuilt rather than ported.
Install
python -m pip install --upgrade --pre 'connectonion==1.8.8b7'
co --version
co benchmark --help
ConnectOnion