Benchmarks
We publish our numbers, including the ones that make a release look worse. A benchmark that only appears when it flatters the model is marketing, not measurement.
On the v1.3 general-knowledge regression
BOSS Offline AI v1.3 scores 16.7% on our general-knowledge evaluation, down from 40% in v1.1. This is a real regression, and it is published deliberately. v1.3 was tuned for instruction-following and structured output, and the trade-off made it worse at broad factual recall. Choose the version that matches your task: v1.3 for reasoning-heavy and structured work, v1.1 for general knowledge. We would rather tell you this than quietly drop the metric.
BOSS Offline AI v1.1
2025-11| Metric | Result |
|---|---|
| General knowledge | 40% |
BOSS Offline AI v1.3
2026-02| Metric | Result |
|---|---|
| General knowledge | 16.7% |
Not yet measured
We are not publishing figures for reasoning, coding, latency, memory footprint, or hardware, because those numbers are not finalised. An honest "not yet measured" is worth more than a number we cannot reproduce. These will appear here as each evaluation is completed.
Reproducibility
Our goal is that anyone can re-run our evaluations and get the same results. The evaluation harness is in progress and will be published alongside these numbers; until it is, treat the figures above as reported-but-not-yet-reproducible by third parties. We will not claim reproducibility we have not enabled.
Placeholder: harness link, hardware specification, and full per-metric tables are pending from the COO once the evaluation harness is released.
See the release notes on the changelog and the artefacts on the projects page.