All plugins

Benchmarks

5 skills

Single-thread, project-scoped, and machine-wide LLM cost, effort, and timing reports across local coding harnesses.

View this plugin on GitHub
any skills-format agent
$ npx skills add rj11io/11ai --skill 11ai-benchmarks-faq --skill 11ai-benchmarks-machine --skill 11ai-benchmarks-pricing-update --skill 11ai-benchmarks-project --skill 11ai-benchmarks-single-thread
claude code
$ claude plugin install 11ai-benchmarks@11ai
codex
$ codex plugin add 11ai-benchmarks@11ai

First time? Add the marketplace once with claude plugin marketplace add rj11io/11ai or codex plugin marketplace add rj11io/11ai, then install any plugin from it.

11ai-benchmarks-faq

Answer questions about the 11ai-benchmarks plugin: which reporting skill fits a job, how the analyzers discover and classify harness usage, how pricing, timing, and cost states are derived, what each report section and warning means, and how the skills are maintained. Routes every question to the plugin's own contracts, references, scripts, and tests and answers with a citation. Use when a user asks how the benchmark skills behave, why a report shows a value, which benchmarks skill to run, or what a report section means.

11ai-benchmarks-machine

Inspect all readable Codex, Claude Code, Claude Cowork, Gemini CLI, Cline, Roo Code, and OpenCode usage stores across the machine, plus exported usage files or folders from unsupported harnesses included on request; detect unavailable remote Cowork sessions; classify harness surfaces and billing modes; normalize token counters; calculate attributable USD costs; measure wall and estimated active time; and write Markdown and standalone HTML reports. Use for global LLM spend, token usage, model and effort cost, thread timing, harness coverage, cross-project analysis, or machine-wide reports over exported usage.

11ai-benchmarks-pricing-update

Refresh and preserve the time-versioned token-pricing history used by the 11ai LLM cost reports from official AI-lab sources, synchronize the project, global, and single-thread copies, and validate their schemas, rate periods, resolver, and equality. Use when model prices, discounts, effective dates, aliases, providers, or pricing caveats have changed, when historical-rate backfill is requested, or when adding pricing support for another AI lab.

11ai-benchmarks-project

Inspect a repository plus its project-attributed Codex, Claude Code, Claude Cowork, Gemini CLI, Cline, Roo Code, and OpenCode records; detect unavailable remote Cowork sessions; classify harness surfaces and billing modes; normalize provider token counters; calculate attributable USD costs; measure wall and estimated active time; and write matching timestamped Markdown and HTML reports. Use for project LLM spend, token usage, model and effort cost, thread timing, harness coverage, or recursive cost analysis.

11ai-benchmarks-single-thread

Inspect one project-attributed Codex, Claude Code, Claude Cowork, Gemini CLI, Cline, Roo Code, OpenCode, or exported usage thread plus recursively spawned Codex sub-agent threads; detect unavailable remote Cowork sessions; classify its surface and billing mode; normalize token counters; calculate attributable USD cost; measure wall and estimated active time; and write matching timestamped reports. Use for the cost, tokens, effort, timing, provenance, or harness coverage of one exact thread tree.