Skip to content
Benchmarks

Benchmarks

We publish our numbers, including the ones that make a release look worse. A benchmark that only appears when it flatters the model is marketing, not measurement.

On the v1.3 general-knowledge regression

BOSS Offline AI v1.3 scores 16.7% on our general-knowledge evaluation, down from 40% in v1.1. This is a real regression, and it is published deliberately. v1.3 was tuned for instruction-following and structured output, and the trade-off made it worse at broad factual recall. Choose the version that matches your task: v1.3 for reasoning-heavy and structured work, v1.1 for general knowledge. We would rather tell you this than quietly drop the metric.

BOSS Offline AI v1.1

2025-11
MetricResult
General knowledge40%

BOSS Offline AI v1.3

2026-02
MetricResult
General knowledge16.7%

Not yet measured

We are not publishing figures for reasoning, coding, latency, memory footprint, or hardware, because those numbers are not finalised. An honest "not yet measured" is worth more than a number we cannot reproduce. These will appear here as each evaluation is completed.

Reproducibility

Our goal is that anyone can re-run our evaluations and get the same results. The evaluation harness is in progress and will be published alongside these numbers; until it is, treat the figures above as reported-but-not-yet-reproducible by third parties. We will not claim reproducibility we have not enabled.

Placeholder: harness link, hardware specification, and full per-metric tables are pending from the COO once the evaluation harness is released.

See the release notes on the changelog and the artefacts on the projects page.