Our drafting agent made a mistake during an internal test on Asthra, our agentic platform for regulatory writing. It took results from a subgroup table and presented them as the overall safety results.

Our automated evaluator missed the population error and gave the section a 100% quality score.

The numbers were real. You could find them in the source. They just didn't answer the question the paragraph was supposed to answer.

I've written before about why solving hallucinations doesn't finish the job. This was a fairly uncomfortable example from our own work.

We were testing changes to Asthra's drafting instructions on two sections: a brief adverse-event summary and the primary efficacy analysis. We tested this using an actual clinical study report and its supporting source documents, all publicly available. The published clinical study report was kept out of the drafting library and used separately as our reference for evaluation.

This was a small internal test. Two sections, one study. But it exposed something that a bigger scorecard could easily have hidden.

The table was genuine. The population was wrong.

The safety draft reported counts and percentages for adverse events, related events, serious events and withdrawals across the treatment groups. The paragraph looked quite ordinary.

When we opened the original PDF, the heading above that table made the problem clear. It described a particular subgroup of participants. The draft had presented those results as though they covered the overall safety population.

The counts and percentages matched the subgroup table. The overall results were in a different table, with different denominators. Both tables contained real results. Only one answered the question we had asked.

A check that asks whether a count and percentage appear in a source can pass that sentence. Even a link to the exact table doesn't settle it. The reviewer still needs to know who is in the table and whether that is who the sentence is talking about.

There was a reviewer flag asking for a denominator tie-out. But the paragraph had already presented the subgroup results as the study-period results. Leaving a question at the end didn't undo the claim.

The original footnotes exposed another error. The draft described a shorter adverse-event observation window than the source specified. Those details are easy to lose when a table becomes extracted text and a summary of that text becomes prose.

A real number can still support the wrong claim. Check the population, observation period, and analysis, as well as whether the value appears in a source.

What the two drafts taught us to check. These are findings from one internal test, not a performance benchmark.

Our first assessment needed correcting too.

We initially reviewed the drafts against the source passages saved in the retrieval trace. In one safety table, the extracted participant-count cell for serious adverse events was truncated. The draft nevertheless reported a count.

That was a legitimate question about the evidence the drafting system had used. It wasn't enough to call the number wrong.

The original CSR showed the intact cell. The draft's number was correct. We corrected the evaluation.

I think this distinction matters if you're building one of these systems. An evaluator saying “unsupported” is a finding to investigate. It isn't automatically a factual error. We have to be willing to check the evaluator as carefully as the draft, including when it appears to confirm a problem we already suspected.

The efficacy draft had almost the opposite problem.

All the reported treatment contrasts matched the original CSR. So did their confidence intervals and unadjusted p-values. Several p-values that the automated evaluator had flagged were actually present in the source.

But the draft couldn't establish whether the trial had met its primary success criterion. It hadn't retrieved the supporting multiplicity-controlled results, so it left a gap.

The original report clearly stated that the primary success criterion had been met. It included the adjusted results and the statistical inference table supporting that conclusion.

I would much rather see a gap than an invented success claim. Still, the writer had asked for the primary efficacy result, and the system had left them to recover the main conclusion from documents where it was available. That's unfinished work, even if the decision to withhold the conclusion was sensible given what the model had seen.

“Not found in the passages retrieved” and “not in the source documents” are very different statements. Our product and our evaluations have to preserve that difference.

What this changed about our approach

A single score didn't explain either section very well. The safety draft had a serious population error hiding behind correctly copied values. The efficacy draft had accurate reported values, false alarms from the evaluator, and an important omission.

For a result like this, I want to see the claim next to its evidence, with the population, treatment, time period and analysis identified. I want missing information separated from incorrect information. And I want to know when the source extraction is too damaged to support a judgment.

The practical questions are quite specific. Did the subgroup heading survive retrieval? Did the footnotes stay with the table? Can we recover an intact cell from the original page? Did we retrieve the joint statistical inference alongside the individual endpoint results? A longer instruction saying “be accurate” doesn't answer those questions on its own.

This test changed how we approach source validation in Asthra. When a number is correct, the next question is: correct for which patients, at which time, under which analysis? That's where this draft went wrong. That's where the evaluation needed to look.

More on the engineering changes in a follow-up.


This account describes an internal Asthra test using a publicly available CSR and its supporting source documents. The errors discussed were in our generated drafts and evaluation process. We have omitted study identifiers and exact results; the findings are unchanged.