HOW WE WORK
We don't sell big projects to strangers. Every engagement starts with a fixed-price measurement. If the measurement says you're fine, we shake hands. If it doesn't, the findings, not a salesperson, make the case for what's next.
01
One week · fixed scope
Day 1
We collect 30 real questions your team actually asks, and you provision a read-only replica or masked snapshot. We never touch production, and we never get write access.
Day 2
We verify ground truth for every question with your team. This surfaces the definition fights on its own, and most clients learn something before any agent runs.
Day 3 to 4
We run the benchmark: your agent against your raw setup, then against a documented, governed version, and through your existing tooling if you have any. Every answer logged, scored, classified.
Day 5
You get the report and a live walkthrough: your score, your failure taxonomy, your fix plan. The eval suite transfers to your repo.
The questions your team genuinely asks: revenue by segment, churn last quarter, top accounts by usage, and the policy calls your agents make. Never synthetic puzzles. Thirty questions is enough to expose every failure class without turning the week into a research project.
02
Scoped from your findings
Scoped only from audit findings. The work, in plain terms:
Settle the definitions.
We interview the people who disagree, get one owner to sign off per metric, and write it down. This step is where accuracy actually comes from; everything after is careful transcription.
Clean the models.
Staged, documented, tested dbt models with explicit grain, the boring engineering that decides whether answers can be right.
Encode the metrics and policies.
Every measure, dimension, and access rule defined once in open-source, version-controlled code (Cube or dbt). The agent stops guessing what a number means, because it's no longer allowed to guess.
Control what the agent sees.
Read-only access where row- and tenant-level rules are compiled into every answer, so an agent acting for one customer cannot retrieve another customer's data.
Prove it.
The audit benchmark re-runs. The engagement ends with two numbers side by side, and the second one is yours to publish internally.
Hand it over.
Documentation, training, walk-away rights.
03
Monthly · cancel anytime
Accuracy is a state, not an event. Systems drift. Teams invent metrics. Model providers ship updates that change behavior overnight. The retainer keeps the number honest: new metrics and policies modeled monthly, re-benchmarks on model releases, drift alerts on your dashboards, quarterly reviews. Cancel anytime, and everything already lives in your repo.
| The Foundation | Hire a data engineer | Enterprise software | |
|---|---|---|---|
| Cost shape | One-time, fixed scope | Full-time salary, ongoing | Annual license, plus your modeling time |
| Time to trustworthy answers | Weeks | 3 to 6 months to hire, then the work starts | Months of internal modeling |
| Who owns the result | You, open-source code in your repo | You | The vendor |
| Proof it worked | Before/after benchmark score | None | None |
One week. Thirty questions. A number instead of a hope.