RESEARCH & FIELD NOTES

What we measure, published.

Benchmarks, failure taxonomies, and build notes from making AI answers trustworthy on open-source stacks. No jargon fog, receipts included.

In progress

  • We benchmarked AI's SQL accuracy on a deliberately messy database.
  • A field guide to plausible wrong answers.
  • What actually changed: raw score to fixed score.
  • The open-source accuracy stack: dbt + Cube + governed access on Postgres and ClickHouse.
  • The AI Data Accuracy Eval Kit is open source.

Find out what your AI actually scores.

One week. Thirty questions. A number instead of a hope.

Get your accuracy scoreOne week · fixed scopeYou keep the eval suite either way.