Teaching a data agent the rules of the data.
A life-sciences team wanted researchers to ask questions about real-world oncology data in plain English and get patient counts back. The agent worked, mostly. The trouble was that its wrong answers looked exactly like its right ones.
Pick a question, then change what the agent is given. These are the real results from our test on synthetic data.
{{av.why}}
Each licensed data product comes with a data dictionary and a usage manual. We read both and turned them into short, single-statement rules, each pointing back to its page.
Before a person approves a rule, it is run against the real tables. Do the columns exist? Do the values match what is stored? Is it really one row per patient? The manual and the dictionary disagreed on values like yes and True, so this step mattered.
For each question the agent gets the approved rules and the real column values for the tables it needs, and nothing else. A query that returns zero rows goes back as a failure, not as an answer.
The agent could always write SQL. What it lacked was the fine print of the data. Once it had that, 16 of 16 test questions came back right, with no retries.