# 1.8.8b7 — write the standard before the skill

An opt-in preview. Stable remains 1.8.7, and the full Personal Wiki acceptance
target remains 1.9.0.

## Skill benchmarks

A skill is a prompt, and a prompt edited without a fixed test gets better on
the example you are looking at and worse on the three you are not. So the
test comes first:

```sh
co benchmark --help                     # the schema, an example, the next steps
co benchmark check reimbursement        # .co/benchmarks/reimbursement.yaml; never runs an Agent
co eval run reimbursement --agent agent.py --skill reimbursement --runs 3
co eval report reimbursement --latest
```

A benchmark is at least five cases, normal and counterexample, each an input
with outcomes that `must` and `must_not` happen. It names no agent and no
skill, so the same cases compare two of them. `co eval run` imports your real
`agent.py` — the `co create` template included, without starting its server —
and scores every expectation PASS, FAIL or UNVERIFIED with the reason and the
evidence:

- a forbidden outcome that happened is a hard FAIL;
- an outside effect the Agent only *claims* — "submitted", "sent" — with no
  tool result showing it is UNVERIFIED, never PASS;
- with `--skill`, a case where the skill did not run fails, however good the
  answer was.

Every run is kept as a read-only report, and `co eval report` compares it with
the run before: score, the cases that started or stopped passing, forbidden
outcomes that came back, and a warning when anything but the skill changed.

![A skill edit scored against the same five cases](https://raw.githubusercontent.com/openonion/connectonion/v1.8.8b7/docs/releases/assets/v1.8.8b7/benchmark-report-after-a-skill-edit.png)

That report is from a real model. The same reimbursement benchmark ran three
times. On the first run the Agent never called the skill: 0/5, although every
answer was right. On the second, a naive skill ran explicitly and scored 3/5,
with three forbidden outcomes. On the third, where only `SKILL.md` had changed,
it scored 5/5. See [docs/cli/benchmark.md](https://github.com/openonion/connectonion/blob/v1.8.8b7/docs/cli/benchmark.md).

The older `co eval [name]` over `.co/evals/*.yaml` works exactly as before.

## What a page sent, and its cookies

`co browser` now has the network layer to go with the DOM, using the tab
names `-t` already has:

```sh
co browser -t shop network requests --type xhr,fetch --status 4xx
co browser -t shop network request 7
co browser -t shop network har start
co browser -t shop network har stop      # ~/.co/browser/har/shop-<time>.har
co browser -t shop cookies               # the cookies of the tab's current site
co browser -t shop cookies save          # ~/.co/browser/cookies/shop.json
```

A HAR opens in Chrome DevTools or Charles, and Playwright's
`route_from_har()` replays it with the site switched off. Header and cookie
values are shaped (`x-sign: <32 hex>`, `<22 chars>`) unless you pass `--raw`,
in the listing and in the file alike. HAR and cookie files are written
owner-only. Flags now take `--key value` as well as `--key=value`. See
[docs/co-browser.md](https://github.com/openonion/connectonion/blob/v1.8.8b7/docs/co-browser.md).

## Finished work that had not reached main

Before cutting the next stable we went looking for work that was finished but
stranded: adapters written against a package that no longer exists, half of a
CLI whose other half had shipped, a July branch that never became a PR. It is
all on today's code now. None of the new commands has run against the real
service yet, and `co --help` says so.

![co --help marks the new commands experimental](https://raw.githubusercontent.com/openonion/connectonion/v1.8.8b7/docs/releases/assets/v1.8.8b7/help-marks-experimental.png)

### Discord and Telegram as inboxes

`co discord` and `co telegram` now work like `co feishu` and `co whatsapp`:
a listener writes each message to disk, and `receive`, `reply`, `done` and
`consume` take it from there.

```sh
co discord listen                 # Gateway connection; DISCORD_BOT_TOKEN in keys.env
co discord receive -t 60
co discord reply <id> "on it"

co telegram listen                # long polling; TELEGRAM_BOT_TOKEN
co telegram receive
co telegram send <chat> "hello"   # unchanged since 1.7.0
```

A wrong token, a Discord bot without the Message Content intent, or a Telegram
bot that already has a webhook stops `listen` with exit 3 and says which,
instead of retrying forever. A Discord reply carries a nonce, so a retry does
not post it twice, and it cannot ping `@everyone`. Edit, delete and react are
not wired for either and say so. Experimental: tested against fakes, not yet a
live bot. See [docs/cli/discord.md](https://github.com/openonion/connectonion/blob/v1.8.8b7/docs/cli/discord.md) and
[docs/cli/telegram.md](https://github.com/openonion/connectonion/blob/v1.8.8b7/docs/cli/telegram.md).

### TikTok post plans

`co tiktok post video.mp4 --account @you` writes a local plan sealed to the
file, the caption and the handle; `--confirm <digest>` checks the seal and then
refuses, because nothing uploads yet. `co tiktok inspect --tab <name>` saves
evidence from a tab you already own and reports `login_required` or what it
could not verify. Experimental. See [docs/cli/tiktok.md](https://github.com/openonion/connectonion/blob/v1.8.8b7/docs/cli/tiktok.md).

### Gemini image output

Gemini image models (`gemini-*-image*`) return their pictures on
`LLMResponse.images` as data URLs; an Agent keeps them on `agent.last_images`
and sends each to a connected client; a new tool writes one to a file:

```python
from connectonion import generate_image
generate_image("a lighthouse at dusk", save_to="lighthouse.png")
```

Direct Gemini keys and OpenRouter only. `co/` models are refused before any
request, because that route has not been shown to return images. See
[docs/useful_tools/generate_image.md](https://github.com/openonion/connectonion/blob/v1.8.8b7/docs/useful_tools/generate_image.md).

## Fixed

- `Agent(tools=[skill])`, the documented way to let an Agent choose a skill,
  raised `NameError` at construction. It builds now.
- An Outlook send that met a 504 was retried, and a gateway timeout usually
  means the mail already went: the same email could reach someone twice. Only
  reads are retried on 503 and 504 now; a 429 is still retried for anything.
- `co wiki` and `co claude` were listed beside stable commands with nothing to
  tell them apart. `co --help` marks them experimental, with `co discord` and
  `co tiktok`; `co telegram` names its shipped `send` apart from the new verbs.
- An Agent whose model answered with only a picture raised "LLM returned an
  empty terminal response". It returns "Generated 1 image." now.

## Known limits

- A plain `Agent` with the skills plugin is not yet told which skills exist,
  so `--invoke auto` cannot pass on one (#1666). `co ai` agents are unaffected.
- `--live` sets `CO_EVAL_LIVE` for tools and skills to read. It does not by
  itself stop a tool that ignores the variable, so run benchmarks against
  fixtures or a safe environment.
- Gemini image calls report no tokens, so they add $0 to `total_cost` although
  Google charges for them.

## Not in this preview

- `co whatsapp-cloud` (#1673) is ready but waits on O API's messaging ingress
  (openonion/oo-api#227), which is not deployed; without it the command could
  only exit 3.
- Remote station status was built on a layout that no longer exists and will be
  rebuilt rather than ported.

## Install

```sh
python -m pip install --upgrade --pre 'connectonion==1.8.8b7'
co --version
co benchmark --help
```
