All questions
Data leadersInvestors

How do you measure and prove accuracy over time?

Three layers: refuse to guess at question time, record every model call in an append-only ledger, and score releases against golden question sets in a standalone evaluation service.

At question time

The system is built to decline rather than improvise. When a question does not map cleanly onto the schema it asks, with candidate columns drawn from the live schema. An answer that arrives without a clarification is one where the mapping was unambiguous.

Continuously

An append-only credit ledger records every model operation with per-agent attribution, latency and outcome. When a number is disputed, there is a trace showing which agent produced it, on what, and how long it took.

Per release

A standalone evaluation service with golden question sets per vertical, six specialized scorers, and regression tracking across seven accuracy dimensions. This is the framework listed on this page as in development, and it is deliberately separate from the product so it cannot grade itself on the runtime's terms.

The honest version

Accuracy claims from a text-to-SQL vendor should be treated as a hypothesis until they are tested on your schema and your vocabulary. That is precisely what the pilot is for: golden questions from your domain, scored against your data, before anything converts.

See it against your own data

A pilot is scoped to one governed use case and time-boxed to eight to twelve weeks, with success criteria agreed before it starts.

Get started free