assaio

skill
Guvenlik Denetimi
Gecti
Health Gecti
  • License — License: Apache-2.0
  • Description — Repository has a description
  • Active repo — Last push 0 days ago
  • Community trust — 76 GitHub stars
Code Gecti
  • Code scan — Scanned 12 files during light audit, no dangerous patterns found
Permissions Gecti
  • Permissions — No dangerous permissions requested

Bu listing icin henuz AI raporu yok.

SUMMARY

Is your AI coding spend delivering? Offline-first analytics for Claude Code, Codex CLI, Gemini CLI, GitHub Copilot CLI and Cline — $/100 AI lines per project, every verdict carrying its own confidence, a self-contained dashboard, exec plugins in any language, and a self-hosted team server. No telemetry; prompts are never read.

README.md

assaio

Is your AI coding spend delivering? assaio shows which projects turn AI budget into
code — and where the same spend would go further — fully offline today, across Claude
Code, Codex, Gemini CLI, Copilot CLI, Cline, and Antigravity CLI. The first piece of a
self-hosted AI-engineering analytics platform.

CI
Go Report Card
Go Reference
OpenSSF Scorecard
License
Latest release

assaio.dev · Roadmap · Backlog · Features · Privacy · Contributing


Your coding assistant already ships stats. Claude Code has
/usage
for plan consumption and /insights for a local 30-day report; Codex CLI has its own
/usage, a year of token
activity read from your OpenAI account; GitHub
Copilot and
Claude Code analytics
publish adoption, accepted lines and cost per commit. They are good, several of them get
better every month, and for one vendor they are often enough.

assaio is for the questions they structurally cannot answer — see
the comparison below. It reads the session logs already
on your machine, across Claude Code, Codex, Gemini CLI, Copilot CLI, Cline and Antigravity
CLI, and keeps them after the tools delete their own. Every figure is computed, not summarized:
same window, same answer, with its provenance, its coverage and its error bars attached. No
account, no upload, about 60 seconds.

assaio-agent effectiveness --by project: AI lines produced, edits, rejections, cost, and $ per 100 AI lines for each project

What the built-in stats cannot do

Every vendor's own numbers get better every month, so a feature-by-feature comparison against
them is worth about one release. These four limits are not a feature gap, and no release
closes them:

  • No vendor will put a competitor's cost on the same axis as its own. The question a
    budget actually turns on is which tool produced more per dollar — and that question only
    exists on an axis nobody selling one of the tools has a reason to draw. assaio prices
    every source it reads against one table and prints them in one report.
  • No vendor will advise you to spend less with it. subscription-fit weighs your
    window's API-equivalent against the plan you actually pay for, and model-right-sizing
    names the premium turns a cheaper model might have handled. Both are only worth reading
    from something with nothing to sell you.
  • A vendor's history is the vendor's to end. Claude Code deletes its transcripts on its
    own schedule and every reader of them goes blind at the same moment — on the machine below,
    22 of 52 days now exist nowhere but this store. Codex shows a year of activity held on
    OpenAI's side instead, which is more durable and still not yours to set. A store on your own
    disk is the only copy whose retention is your decision.
  • A number a language model produced cannot be trended. /insights summarizes sessions
    with a model, so two runs over one window need not agree — and a figure you cannot reproduce
    is one you cannot compare against last month or defend in a budget meeting.

And the fifth, which is a refusal rather than a limit: assaio will not rank your
engineers.
Vendor analytics reach the individual by design — GitHub's Copilot metrics APIs
publish data at the enterprise, organization, repository and user
levels
.
That is the natural shape of a product sold by the seat, and it is the shape assaio
declines. --by member is refused with a reason rather than caveated, in table, JSON and CSV
alike; the dashboard's team panel shows a member's engagement while output and spend stay the
team's total; and every export renders members as stable pseudonyms unless an operator passes
--identify, which names individuals and says so in its own output. What that operator
does next is theirs — the tool builds them no leaderboard.

That is not a privacy footnote. For a whole class of organization it is the deciding property:
where a works council, a co-determination agreement or a local data-protection regime has to
sign off on a measurement tool, it can rank individuals and we have promised not to is a far
harder sentence than it does not. This is not a legal opinion — every agreement and
jurisdiction is its own question — only the answer that conversation asks for.
PRIVACY.md has the mechanism in full;
BACKLOG.md lists it beside the other
things this project will not build regardless of demand.

Then, this month, feature by feature

Measured against one real machine — a Claude Max user's own store, 178,254 records over 52
days — beside what these tools ship. The /insights column describes the command as
documented and as
observed on Claude Code 2.1.238 in August 2026. The Codex column describes /usage as shipped
in Codex CLI 0.140.0, read from
that release's own source and wire contract in September 2026. Both are fast-moving products
and this comparison carries its dates for that reason.

Claude Code /insights Codex /usage assaio
How far back 30 days, and only what the tool still keeps A lifetime token total and 12 months of daily activity, held by your OpenAI account rather than by the log Everything ever ingested. On that machine 22 days are older than Claude Code's own retention — they exist nowhere else now
How much A capped sample of sessions per run Every day the account recorded Every session the store holds
How An LLM summarizes the sessions, so two runs need not agree Counted, not summarized — the same total twice Computed from the records. Same window, same number, every time
Which tools Claude Code Codex, on the account you are signed into Claude Code, Codex CLI, Gemini CLI, Copilot CLI, Cline — normalized, one cost basis. Antigravity CLI joins for sessions and activity, and is excluded from every cost figure rather than counted at zero
What moved A snapshot A 52-week heatmap of one number digest reports the delta since last run, and says when two runs are not comparable
When it was wrong Nothing restates a past report Nothing restates a past report A parser fix reaches stored history, and a re-read that lowers a figure is counted and reported
A team Not aggregated One account serve + sync, pseudonymous by default, nothing ranked per person
Where the line is "Friction", "satisfaction" as scores No verdict, and none claimed A verdict only where its line is derived from your data, cited, or set by you. Fourteen report their figure and refuse the grade

Every correction that row promises is written down. docs/corrections.md is the
register: what was wrong, since when, what the wrong number showed a reader, and which release
put it right — including the times a second review pass overturned the first fix.

What Claude Code's /usage does that no local reader can. It reads your plan consumption
from the API — the 5-hour and weekly percentages, and your extra-usage balance. No local
transcript carries any of that (verified: zero limit-shaped fields), so assaio cannot report
it and does not try. subscription-fit answers a different question — whether the plan beats
the API-equivalent estimate of what you ran — and the two are worth reading together.

And what Codex's /usage does that no local reader can. It draws a year of token activity
— a lifetime total, your peak day, your current and longest streak, your longest single task,
and a 52-week heatmap — from OpenAI's own record of the account rather than from ~/.codex.
That is not bounded by the rollout logs still on disk, so on a laptop set up last week it
reaches back further than assaio possibly can, and it is counted rather than summarized.
Two things are worth stating beside it, because they are the shape of the whole comparison
rather than a knock: one day of that history is a single integer — a token total with no
input/output/cache split, no model, no project and no $ — and it is one account on one
vendor. Codex's plan and rate-limit view lives in /status, and like Claude Code's it comes
from the API. assaio fetches none of it, and does not try.

Privacy first

assaio is built to be safe to run on a work machine without asking anyone's
permission. This section describes the local agent, which is what runs on your machine
and needs no network at all.

  • No network at runtime. The price table is embedded into the binary at build time.
    Nothing is fetched, posted, or phoned home.
  • No telemetry. No usage pings, no analytics, no crash reporting.
  • Prompts and code are never read. The parsers extract token counts, model names,
    timestamps, session IDs, and activity counts (lines added/removed, edits, rejections) —
    never prompt text, never file contents. Line counts come from diff +/- markers;
    the code on those lines is counted, never stored.
  • Your data stays local. Everything lives in one SQLite file under your home
    directory. clear --all --yes deletes it.

Full detail, including the exact fields extracted: PRIVACY.md.

The optional team server (serve + sync) pools a team's usage on infrastructure you
stand up. It and the experimental runtime inspect are the only commands that touch the network, and
only when you invoke them — the guarantee above is about the local analysis. Team views stay
aggregated and pseudonymized by default, and the refusal above holds there too: the per-member
row shows engagement only, output and spend are the team's, and nothing is ranked per named
individual.

Install

assaio-agent is a single static binary — no CGO, no runtime dependencies — built for
macOS, Linux, and Windows (amd64/arm64), and the test suite runs in CI on all three.

Homebrew (macOS, and Linux via Linuxbrew):

brew install assaio/tap/assaio-agent

Any platform, with Go 1.25+ (the release toolchain is 1.26.6, the patch level govulncheck requires):

go install github.com/assaio/assaio/cmd/assaio-agent@latest

Or take your platform's archive straight from
Releases.

Manual install recipes — Linux / macOS and Windows
VER=$(curl -fsSL https://api.github.com/repos/assaio/assaio/releases/latest |
  sed -n 's/.*"tag_name": *"v\([^"]*\)".*/\1/p')
curl -LO "https://github.com/assaio/assaio/releases/download/v$VER/assaio_${VER}_linux_amd64.tar.gz"
tar xzf "assaio_${VER}_linux_amd64.tar.gz" assaio-agent
sudo install assaio-agent /usr/local/bin/
$ver = (Invoke-RestMethod https://api.github.com/repos/assaio/assaio/releases/latest).tag_name.TrimStart('v')
Invoke-WebRequest "https://github.com/assaio/assaio/releases/download/v$ver/assaio_${ver}_windows_amd64.zip" -OutFile assaio.zip
Expand-Archive assaio.zip -DestinationPath "$env:LOCALAPPDATA\assaio"
[Environment]::SetEnvironmentVariable("Path", "$([Environment]::GetEnvironmentVariable('Path','User'));$env:LOCALAPPDATA\assaio", "User")

New terminal after the PATH change; on ARM replace amd64 with arm64. Scoop and winget
packages are on the backlog.

Every release artifact ships with checksums, an SPDX SBOM and a build-provenance attestation —
verify one with gh attestation verify <archive> -o assaio.

New here? assaio-agent demo prints the full reports on bundled sample data — no logs
needed — so you can see the value before importing your own history.

Quick start

The commands below are identical on macOS, Linux, and Windows (PowerShell or cmd) —
each tool's log location is auto-detected per OS, and assaio-agent doctor shows
exactly what was found on your machine.

Import your local history, then report on it. The first backfill reads every session
log your AI coding tools have written — often months of data.

$ assaio-agent backfill
claude-code   files=3  records=4  inserted=4  steps=9
codex         files=1  records=1  inserted=1  steps=2
gemini-cli    files=0  records=0  inserted=0
copilot-cli   files=0  records=0  inserted=0
cline         files=0  records=0  inserted=0
agy           files=0  records=0  inserted=0

First, where the money goes — spend per project:

$ assaio-agent report --since 30d --by project
+----------+--------+--------+-----------+---------+--------+--------+
| PROJECT  |     IN |    OUT |   CACHE R | CACHE W | CACHE% | COST $ |
+----------+--------+--------+-----------+---------+--------+--------+
| api      | 22,680 | 16,800 | 1,664,000 |  38,000 |   98.7 |   1.56 |
| planning |    270 |  7,300 |   430,000 |  24,000 |   99.9 |   0.55 |
| webapp   |  1,720 | 15,000 |   848,000 |  80,000 |   99.8 |   1.31 |
+----------+--------+--------+-----------+---------+--------+--------+
|          |        |        |           |         |  TOTAL |   3.41 |
+----------+--------+--------+-----------+---------+--------+--------+

But spend is only half the question. Which projects turn that spend into code, and
which don't?
That is effectiveness — AI output over cost, per project:

$ assaio-agent effectiveness --since 30d --by project
+----------+----------+-------+-----+--------+-------------+
| PROJECT  | AI LINES | EDITS | REJ | COST $ | $/100 LINES |
+----------+----------+-------+-----+--------+-------------+
| api      |       22 |     1 |   1 |   1.56 |        7.08 |
| planning |        0 |     0 |   0 |   0.55 |           — |
| webapp   |      336 |     2 |   0 |   1.31 |        0.39 |
+----------+----------+-------+-----+--------+-------------+
| TOTAL    |      358 |     3 |   1 |   3.41 |        0.95 |
+----------+----------+-------+-----+--------+-------------+
Efficiency is directional: task type (greenfield vs. debugging) drives lines-per-cost; this is a diagnostic per project, never a performance metric.
Not every source records changed lines; the ones that do not contribute cost but no line counts -- run `assaio-agent signals coverage` for what your own data supports.
Cost is an estimate at public pay-as-you-go API prices -- not your actual spend; subscription plans bill a flat rate and differ.

The headline column is $/100 LINES — cost per 100 AI-written lines. Read this table:

  • webapp turns spend into code: greenfield work, 336 AI lines at $0.39 per 100
    lines
    .
  • api costs $7.08 per 100 lines — 18× more. It is debugging-heavy: many
    tokens spent reading and reasoning, few new lines — so the ratio reads worse.
  • planning shows : real cost, zero lines. That is an architecture session,
    not waste — so assaio prints an honest blank, never a fake $0 or a divide-by-zero.

This is a per-project diagnostic, not a scoreboard. The ratio is directional: task
type drives lines-per-cost far more than any person does, which is why assaio groups by
project here, not by author, by default.

Group by day, project, tool, model, or entrypoint. Want machine-readable
output? Add --format json or --format csv. Not sure what was detected? Run
assaio-agent doctor.

For the fuller read, assaio-agent analyze prints a short directional report for each
metric — adoption, model fit, context health, throughput, rework — and
assaio-agent dashboard writes a self-contained, offline HTML dashboard you can open in a
browser or hand to a teammate (project names pseudonymized by default). All of it runs
locally, in about 60 seconds.

Cost honesty and control

Every $ assaio prints is an estimate at public pay-as-you-go API prices — not your
actual spend. Token counts are computed server-side, and a flat-rate subscription (Claude
Pro/Max, ChatGPT Plus/Pro) makes the effective cost-per-token entirely different, so
assaio labels every cost figure as an estimate. If you pay a subscription or a negotiated
rate, set your real basis in config.pricing (an effective $/token or a monthly plan
cost) and reports show a truer figure alongside the estimate.

Two commands build on that:

  • assaio-agent check --max-tokens N (or --max-cost N) is an exit-code budget gate for
    CI or a pre-push hook — non-zero when usage exceeds the budget. Token budgets are the
    plan-independent default; a $ budget is allowed but labeled API-equivalent.
  • assaio-agent report --compare (and effectiveness --compare) shows period-over-period
    top movers — which projects' cost and AI lines rose or fell vs. the previous equal
    window.

What comes next is in ROADMAP.md.

What assaio measures — and what it doesn't (yet)

The honest scope. assaio measures how much AI is producing, how efficiently, and with
how much friction
— not how good the result is. That line is deliberate; blurring it
would be the easiest way to lie with this tool.

Measured today — fully offline, per project / model / tool:

Dimension What it is How it's derived
Adoption Tool mix, sessions, and sub-agent delegation. Session logs; sub-agent token usage is now counted (it used to be invisible — a correctness fix).
Effect AI lines added and removed. Counted from the +/- markers in diff hunks. The code text itself is never read or stored — only the line counts.
Efficiency $ per 100 AI lines — a directional diagnostic, never the headline. Priced cost ÷ AI lines added, shown per 100 lines. Unpriced models stay an honest blank.
Friction Edits, and rejections. Edit/Write tool-calls, plus proposals the human declined (REJ).

Not measured yet — these need correlation with your git history and issue tracker, the
deeper server work still ahead (the team-server MVP that ships today pools
usage; it does not yet reach into your repos):

  • Whether AI-written code survived in your main branch after review, rewrites, and reverts.
  • Whether it caused bugs, compared only against age-matched human code.
  • Code quality or maintainability.

So today's answer is "how much is AI producing, how efficiently, and with how much
friction"
— a per-project diagnostic. "Did it actually work, and was it worth it in
quality terms"
is the roadmap.

One more limit is enforced rather than documented: a figure is computed only over the
sources that record its field.
A tool that never writes an edit count is absent from the
session mix rather than counted as a window full of conversations, the reach is stated as the
verdict's signal coverage, and a figure nothing in your window can answer prints and
withholds its verdict (ADR 0011).
assaio-agent signals coverage reads your own mix and says what it supports.

Commands

The ones a first week actually needs: demo to see it on sample data, backfill to import
your history, effectiveness for the headline, analyze for the fuller read, dashboard
for something to hand a teammate, and doctor when a number looks wrong.

Every command
Command What it does
demo Print the full reports on bundled sample data — no logs needed, the 60-second first look.
init First run: show what will be read, import it, write the report, name what to run next.
backfill Import all historical local session logs into the store.
report Print a token/cost report. --since 7d, --by day|project|tool|model|entrypoint|task|outcome|difficulty, --format table|json|csv, --compare for period-over-period top movers.
share Render one window as something postable: a square reel, a still poster, the post text — or --format text for a terminal block. Redaction is structural, not a flag: repositories are counted, never named. Every figure is quoted from analyze; every frame carries its measurement layer. Imports on an empty store, so a fresh install reaches a card in one command.
effectiveness Print AI output vs. cost — AI lines, edits, rejections, and $/100 AI lines — per project. Same --since, --by, --format, --compare flags (defaults to --by project). A directional, per-project diagnostic.
analyze Run metric validators — adoption, model fit, context health, throughput, rework, plus any configured metric plugins — and print each one's directional report, led by the few findings worth a week's attention with the reasons that ordered them. A window with nothing worth acting on says so instead of promoting the least weak read. --since, --format text|json, --list, or pass [name...] to run a subset.
check Exit non-zero when usage exceeds a budget — --max-tokens N (plan-independent default) or --max-cost N (labeled API-equivalent) — or when a configured rule plugin raises an error alert. A CI / pre-push gate.
dashboard Write a self-contained, offline HTML dashboard — stat tiles, hot/going-stale projects, model/tool mix, inventory. --since, --output. Project names are pseudonymized by default so it's safe to share; --no-anonymize for real names.
serve Run the self-hosted team server: collects usage pushed by teammates' sync and serves the aggregated, pseudonymized-by-default team dashboard.
sync Push this machine's local usage to a team server — pseudonymous by default, --member is an explicit opt-in to a real name.
recommend The few experiments this window's evidence supports, each as a typed record: what triggered it, what it requires, what it risks, how to undo it, when to look again, and the figure that would show whether it worked. The rendered text projects the record and adds nothing to it. A thin window, a low-confidence verdict or a metric missing a declared input produces nothing at all, and says that is an abstention rather than a clean bill.
reprice Re-price the window already in the store against another entry in the same price table: what this same set of turns would have cost on a different model, and how its projected monthly rate stands against a flat plan price you pass in. Arithmetic over observed events, never a counterfactual — it states what it holds fixed, what share it could not price, and that it claims nothing about another model's output. Ends in one recommend experiment with its rollback and follow-up. --since, --against <model>, --plan "name=monthly-price" (both repeatable), --format text|json. assaio vendors no plan catalogue: a published price changes without notice, so the candidate is a figure you read off your vendor's page.
runtime Experimental. runtime inspect snapshots a self-hosted vLLM server's or NVIDIA DCGM exporter's own metrics endpoint — by URL or from a saved file. Read-only: nothing stored, no cost model, no GPU advice. A counter is never shown as a rate and a missing metric is never shown as zero. A hosted vendor's accelerators stay unknown; no local signal reveals them. May be removed — see ROADMAP.md.
doctor Show detected tools, log locations, store inventory and size, format-drift canaries, how much of your store the price table cannot cost, and accuracy caveats. --strict exits non-zero for cron/CI — including when too much of the store carries no model price for $ to mean anything (pricing.max_unpriced_share, default 5%).
survival Read the local git history beside your AI usage: how much of a repository's recent work still lives in HEAD. Directional and age-dependent by construction — a short window reads near 100% because its commits have had no time to be rewritten — so it is a lead, never a productivity figure. --since, --repo.
status A terminal overview: inventory, headline $/100 lines, hottest projects, and what's going stale — projects only. --since.
statusline Print one ambient line for an editor or shell status bar: today's tokens, AI lines where a source records them, cost basis, and how fresh the data is. The day is your machine's local day. Read-only, and never fails loudly — see automation.
explain Print a metric's long-form page — what it measures, how to read it, what to do about it, and its limits. Needs no store, so it works before your first import; no argument lists every metric.
mark Label a session with what the work actually was — task class, outcome, difficulty. Category values only, never free text, and never sent by sync. Defaults to the newest session in the repository you are standing in; --last, an id prefix, --list, --unmark. Every metric can then be read per kind of work. --suggest derives a label from what the store already recorded — branch, skill, sub-agent, entrypoint — and shows the evidence for each; --accept-suggested writes them, never replacing one made by hand. The derivation is a rule engine: add your own convention under labels.rules, and a repository with no convention yields nothing rather than a guess.
digest Write markdown saying what moved since the last digest — totals, per-model and per-project movers, verdict changes — fit for cron or launchd, with delivery left to your own script. States when the comparison is weak: overlapping windows, windows of different lengths, or a parser that changed between the runs and therefore corrected history underneath the numbers. --weekly, --since, --dry-run.
reconcile Compare a vendor's own billing or usage export — a CSV or JSON you downloaded, no credential and no network — against assaio's estimate. Computes the scope mismatch first, names only the parts of the delta that have evidence, and reports the rest as unexplained. Nothing is ever adjusted to make the two sides agree. --since, --map to bind columns, --format text|json. See reconciling.
signals list what assaio can report; describe <id> for what one signal counts, where it is honest, and what a zero means; coverage reads your own store and says which signals your data actually supports, fully, partly, or not at all.
clear Delete stored data — needs an explicit scope (--all, --older-than, --tool, --labels) and --yes. Session labels survive every scope but --labels: no re-import can rebuild them.
compact Reclaim disk space the store freed but still holds — deleting rows alone never shrinks the file.
config Print the effective configuration and where it was loaded from.
plugins list configured exec parser plugins; verify <name> runs one and reports protocol conformance without storing.
metrics list configured exec metric plugins; verify <name> runs one on your real window and reports contract conformance plus the rendered result — nothing stored.
docs export writes everything this binary can enumerate about itself — signals, sources with their depth, validators with their scope, the whole command tree with flags and defaults, every config key with its environment variable, and both halves of the metric contract — as JSON, or as a self-contained HTML page. The published reference is generated from it, and a test fails when the website or the docs and the binary disagree.
version Print the version (also --version).

Configuration

Configuration is optional — the defaults (since: 30d, format: table) are sensible.
To override, create ~/.config/assaio/config.yaml (XDG-aware):

since: 30d
format: table

Environment variables take precedence over the file, using an ASSAIO_ prefix:
ASSAIO_SINCE=7d, ASSAIO_FORMAT=json. Command-line flags win over both. See
config.example.yaml for a documented starting point.

Supported tools and accuracy

Today assaio reads six sources:

  • Claude Code — session transcripts under ~/.claude/projects/**/*.jsonl.
  • OpenAI Codex CLI — rollout logs under ~/.codex/sessions/** and
    ~/.codex/archived_sessions/**.
  • Gemini CLI — chat logs under ~/.gemini/tmp/<hash>/chats/session-*.jsonl.
  • GitHub Copilot CLI — session events under $COPILOT_HOME, else
    ~/.copilot/session-state/<id>/events.jsonl. Totalled when a session ends, so its
    records are session-granularity.
  • Cline — task data under the VS Code extension's global storage
    (saoudrizwan.claude-dev) and the Cline CLI's ~/.cline/data/tasks.
  • Antigravity CLI (agy) — conversation transcripts under
    ~/.gemini/antigravity-cli/brain/<id>/.system_generated/logs/transcript.jsonl. The one
    source with no token counter anywhere in its format: it contributes sessions, turns,
    tool calls and edits, and every token and cost figure withholds its verdict for it rather
    than counting a zero. Its format records no working directory either, so its sessions
    carry no project and are absent from every per-project figure — report --by project,
    effectiveness, and the dashboard's project panels alike. In a tool whose unit of analysis
    is the repository, that is the larger of its two absences. Verified against Antigravity CLI
    1.1.23; ~/.gemini is shared with Gemini CLI and the two are read as separate sources.

Costs are computed from a vendored snapshot of the
LiteLLM price table, and doctor reports how much of
your store that table cannot price rather than leaving you to find out from a figure that
looks complete. We are honest about what we can and cannot measure — assaio-agent doctor
prints these caveats every run, so this list is a preview, not the only place it lives:

The accuracy caveats doctor prints
  • Claude input_tokens can be a streaming placeholder, so totals may diverge slightly
    from the Anthropic Console.
  • Codex reasoning tokens are reported separately but assumed to be included in output
    for cost.
  • Gemini tool-use tokens are folded into output tokens, and ~/.gemini may be shared
    with other tools; the mapping is based on observed samples, pending verification
    against more real traces.
  • Cline stores its own per-request cost, but assaio recomputes cost from tokens for
    cross-tool consistency.
  • The price table carries no long-context premium (e.g. a 1M-context [1m] rate), so cost
    for a very long-context session is an under-estimate. The 5-minute and 1-hour cache-write
    tiers are priced separately where a source records which one a write bought.
  • Days and week-over-week windows are bucketed in UTC, so late local-evening work can
    land on the next UTC day.
  • Activity counts (AI lines, edits, rework) — not tokens or cost — read low on a session
    ingested while it was still being written; the next backfill restates them upward, never
    downward, so a count that first came out too high stays. A parser fix that lowers a stored
    figure is the other direction, and backfill reports those rows as restated-down=.
  • All on-disk log formats are vendor-internal and may change between tool versions.

When a model is missing from the price table, assaio never fakes a $0 cost: the
table shows , JSON reports "cost": null, CSV leaves the cell empty, the row is
excluded from the TOTAL, and the footnote states what share of the window's tokens the
cost could not see. Wrong-but-precise numbers are worse than an honest blank.

More sources — opencode, Aider, Qwen, Roo & Kilo Code — are on the
roadmap, and buildable out-of-tree today.

Adapt it to your organization

assaio is built to be adapted, not just installed. Three out-of-tree surfaces need no fork
and no rebuild, each an executable in any language declared in your config: a parser
plugin
reads another tool's logs, a metric plugin renders beside the built-in verdicts,
and a rule plugin gates check on your own thresholds. In-tree you can add a data source
or a metric validator; outside the binary you can query the documented SQLite schema
directly or pipe --format json|csv into your own tooling. The
extensibility guide has every contract, with a verify command per
protocol.

One honest limit: out-of-tree Go plugins that import assaio as a library are not
possible yet.
The core lives under Go's internal/ on purpose while the API stabilizes.
Exec plugins are the supported path today — their contract is a versioned data format, not a
Go API, so they survive core refactors. An in-process Go API is still ahead; see the
roadmap.

Status and roadmap

assaio is pre-1.0 and already reaches well past an offline token reporter: six parsers,
twenty-one metric validators that each carry a confidence envelope and state which measurement
layer they sit on, format-drift canaries, a published source-depth matrix, offline
reconciliation against a vendor's own export, the offline Assay dashboard, structured
recommendations, and a self-hosted team-server MVP. FEATURES.md is the
maintained inventory, with the release each capability arrived in.

What it still can't do is connect a session to the change it produced — whether AI-written
code reached a pull request, passed review and CI, survived in main, or caused
bugs
. That needs correlating usage with git and pull-request history, and it is a
milestone, not a shipped feature.

Two other things are honestly incomplete and said so on purpose. The team server is an MVP:
authenticated, rate limited and diagnosable as of v0.24, and still without RBAC, TLS of its
own, resumable sync, retention or a tested restore drill. And runtime inspect is an
experiment behind a demand gate — if no self-hosted user finds a recurring decision it
answers, it is removed rather than kept.

ROADMAP.md is the direction, ordered, with each milestone's exit criteria and the
gates that have not been met; docs/compatibility.md is what v1.0
freezes; BACKLOG.md is the ranked pool behind both.

The core is Apache-2.0 and stays that way; a hosted or enterprise offering may fund
development later.

Why "assaio"?

An assay is the metallurgist's test: what tells you how much real metal is in a piece of
ore. Ore can glitter and still be worthless — which is exactly the position engineering
organizations are in with AI coding tools: tokens burned, seats bought, dashboards full of
activity, and very little honest measurement of what any of it is worth. assaio (assay

  • io) is built to be that test, and it would rather show you an honest blank than a
    precise-looking number that is wrong.

The metaphor also sets the ethics: an assay examines the metal, never the miner. assaio
measures value, not people — aggregation and pseudonymization are the defaults in
everything built on this agent.

Contributing

Contributions are welcome. Please read CONTRIBUTING.md first — it is
the authoritative set of rules. In short: sign your commits with the
DCO (git commit -s), one commit per PR
(squash before review), Conventional Commit subjects, and a green CI gate (gofmt,
go vet, golangci-lint, go test -race).

AI assistants get the same rules via AGENTS.md. Adding support for a new AI tool
is a well-scoped first contribution: one parser package, golden-file tests, and a connector
issue template to guide you — the extensibility guide walks it through.

Security

Found a vulnerability? Do not open a public issue — see SECURITY.md for
private disclosure.

License

Apache-2.0 — see LICENSE. assaio embeds a snapshot of the model price table
from the LiteLLM project (MIT); attribution is in NOTICE.

Partner with us

assaio is at the beginning of its roadmap, and the most valuable thing at this stage is
real-world usage.

  • Pilot organizations & design partners. If your team wants to see what its AI coding
    spend produces per project today — and, next, whether that code survives and holds up in
    quality terms — we want to build the next milestone against your reality,
    not our assumptions. Direct influence on the roadmap; no commitment beyond honest feedback.
  • Tool vendors & integrators. If you build an AI coding tool and want it measured
    fairly, help us get your connector right.
  • Anyone running Gemini CLI or Cline. One redacted session log closes the one gap
    nothing here can close alone: both are calibrated against a sample written in the source's
    shape, never a real capture. A redacted vendor billing export is the same kind of gift.
  • Contributors. One tool = one package; one metric = one file. The codebase is
    deliberately small and readable — a good place to make your first open-source
    contribution count.

Talk to us: GitHub Discussions ·
karauda.com/contact · assaio.dev

Yorumlar (0)

Sonuc bulunamadi