Teaching a data agent the rules of the data.
A life-sciences team wanted its researchers to ask plain-English questions of real-world oncology data and get patient counts back. Most of the time the agent got them right. When it was wrong, nothing about the answer said so.
Pick a question, then change what the agent is given. These are the real results from our test on synthetic data.
{{av.why}}
Each licensed data product comes with a data dictionary and a usage manual. We read both and turned them into short, single-statement rules, each pointing back to its page.
Before a person approves a rule, it is run against the real tables. Do the columns exist? Do the values match what is stored? Is it really one row per patient? This step is where the disagreement between the manual and the dictionary over values like yes and True turned up.
For each question the agent gets the approved rules and the real column values for the tables it needs, and nothing else. A query that returns zero rows goes back as a failure, not as an answer.
Writing the SQL was never the problem. The agent did not know how this particular dataset records things: which values mean yes, which dates count, which rows to leave out. Once it did, 16 of 16 test questions came back right, with no retries.