A while ago we made a decision that reads well in a blog post and creates real work in practice. When Asthra drafts a regulated document and can't find a fact in your sources, it's designed to leave a marker rather than fill the gap with a guess. You get [MISSING INFO: batch size] in the draft instead of a plausible number nobody can trace. We've written about this twice, first in end-of-run QC and most fully in Regulatory Review: an FDA audit inside the draft, where a grounded reviewer plants its findings in the document the same way.
Both of those posts stop in the same place. The flag exists, the writer can click to it, and then the writer does the work. That was honest, and it was half an answer. A draft covered in [MISSING INFO: …] markers is far more trustworthy than a draft full of confident invention, but it is still a to-do list somebody works through by hand, and in a real submission that list runs to hundreds of items.
This is the other half. Asthra now resolves the flags it raises. You tell it where the missing information lives, it fills every gap across the submission, and it hands each change back to you as a tracked revision in Word that nobody accepts on your behalf.
Flag resolution, end to end: the fill comes back as a Word tracked revision, and the agent can't turn change-tracking off.
When honest flagging turns into a burn-down list
The scale is easy to underestimate. In one CMC section we measured, a single size-and-shape table carried seventy-one identical [NOT FOUND] markers, the same missing field repeated down every row. Across a full dossier you are looking at hundreds. No writer wants to work through that by hand, and no AI product should hand you the list and call the job finished.
So the question we sat with was narrow. What part of resolving a flag actually needs a person, and what part is transcription nobody enjoys?
What the writer knows that the model doesn't
The answer turned out to be specific. A medical writer usually knows where a missing fact lives. The batch size is in the batch record; the stability data is in the study report. What they don't want to do is open that document, find the number, and copy it into the draft by hand, forty times, across a submission.
So the resolution flow asks the writer for the one thing only they have, the pointer to the source, and does the rest itself. In practice that is a few checkboxes: you pick which library documents Asthra should treat as references for this resolution, and that selection constrains the retrieval planner directly. It is not a hint dropped into a prompt that the model may or may not honor. Scoping the search to three documents means the agent searches those three documents.
Nothing lands in the document unless you accept it
This is the part that matters most, and it is worth being precise about, because it is what separates something you would trust on a submission from something you would not.
When Asthra resolves flags, whether it is one flag, a section's worth, or an entire submission in a single pass, it does not write into your document. It proposes into it. Every edit comes back as a native Word tracked revision: deletions struck through in red, insertions in color, each one attributed, each one accepted or rejected on its own in Word's reviewing pane. You work through them the way you work through a colleague's markup, which is the check-is-a-click bar we've argued regulated AI has to clear.
What makes this more than a checkbox is where the change-tracking lives. Here is the engineering comment, quoted directly, because it says it more plainly than a marketing line would:
"Enable Track Changes before the write so the edit appears as a tracked revision in Word. This is deterministic infrastructure — the agent doesn't call it; it happens automatically for every chat-initiated write. The writer sees red strikethrough (deletions) and colored text (additions) and can accept/reject."
The load-bearing word is deterministic. Turning on change tracking is not a tool the model decides to invoke, and it is not a sentence in a prompt that the model might drift away from under load. It is infrastructure wrapped around every agent write, on a code path the model has no access to. The agent can't edit your document silently because the code never gives it the option.
That distinction is easy to wave past, so it is worth slowing down on. Most AI writing tools promise a human in the loop and implement that promise as an instruction in a system prompt. Instructions like that hold most of the time and fail quietly the rest of the time. Here the writer's approval is not a request made of the model; it is a property of the pipeline the model runs inside.
Two smaller decisions sit next to it. Asthra borrows your Track Changes setting for the length of its own write and then puts it back the way it found it, so if you had tracking off, it goes off again afterward. And acceptance is genuinely per change, not per run. A bulk resolve across forty sections is not one all-or-nothing approval; it is forty sections' worth of individually reviewable revisions, and you can take thirty-eight and reject two.
This is also what makes bulk resolution possible in the first place. We could only build it because of that review step. Rewriting two hundred spans across a submission automatically would be reckless if the result were a silently changed file; with a marked-up draft you review exactly the way you already review every draft, it is just a larger version of accepting a colleague's edits.
Resolving one flag
In day-to-day use, most resolution starts from a single flag. Every section's inspector has an Issues tab, marked with a ⚑ and a live count of what is still open, and there is a submission-wide Issues view with counts across every section, so the whole burn-down is visible in one place.
On any flag you click Resolve with asthra; the tooltip reads "Ask asthra to research and resolve this marker." Asthra switches to the conversation, highlights that flag's paragraph in the live Word document so you can see exactly what it is about to touch, captures it as context, and pre-fills the instruction without sending it. You read what it is going to ask, adjust it if you want, and press send yourself.
That last detail is deliberate. It puts the writer in the loop at the moment of intent, not only at the moment of approval. If you know the number is in the 2024 stability report specifically, you say so before anything runs, instead of correcting a wrong guess afterward. The result comes back as a before/after card with Approve and Reject, and the card already tells you if the target text is ambiguous or has drifted since the flag was raised, so you are warned before you decide rather than after.
Resolving a whole submission
When you want to clear more than one flag, the Resolve issues with asthra wizard runs the same logic at scale, in three steps.
First you choose what to resolve: everything open, a single section, or a hand-picked set, with live counts of open, resolved, dismissed, and stale.
Then comes the step the whole feature turns on, pointing Asthra at the source. The wizard's own words: "Pick which library documents asthra should use as references. Leave all unchecked to search the whole library." This is where your knowledge of where the facts live becomes the input.
Last, a plain statement of what is about to happen, "asthra will re-draft these sections, filling the selected gap / regulatory issues from the library and leaving the rest of each section unchanged," and a Start resolving button. Batches over fifty flags get an explicit warning first.
One engineering choice underneath this is worth surfacing, because it is what keeps a bulk resolve coherent rather than merely fast. The worker groups the selected markers by the paragraph that contains them and rewrites each span once. Those seventy-one identical markers in the size-and-shape table are not seventy-one separate model calls racing each other to fill the same row differently. They are one rewrite of the span that holds all of them, with all seventy-one in view at once. That is the only way the filled-in values come out consistent with each other.
Not every flag is a gap. Some simply do not apply to your product, and for those there is Mark N/A ("Mark this gap as not applicable — no document edit needed"), which records the flag as Done with your note and stays reversible. Dismissed is tracked separately from resolved, so the burn-down never blurs the two: you can always tell what was filled from what did not apply.
Why bulk-editing a submission is safe
Two hundred automated edits across a regulatory document is the kind of thing that should make a regulatory writer nervous, so most of the engineering went into the safety.
Outside the flagged span, the document is preserved byte for byte. The span Asthra sends for rewriting is expanded to the whole enclosing paragraph, deliberately a little wider than the bracket itself, so it can also fix the prose around the marker. If the draft reads "Detail X was not found in the sources" next to a flag, filling the flag and leaving that sentence standing would be worse than not filling it, so the instruction covers that case explicitly. The replacement is a whitespace-tolerant, first-occurrence-only substitution; everything outside that single match is left exactly as it was, which is asserted in code and pinned by a regression test. It is version-checked throughout, so if a section changes while a bulk resolve is in flight, that section is skipped with a message rather than overwritten.
The harder problem is knowing which flag you meant. Identical flags are the norm, not the exception. In the CMC section we measured, 181 of 208 flags, about 87 percent, shared their exact text with another flag in the same section. Finding "the one the writer clicked" is a genuinely hard anchoring problem, and the standard tool for it, the W3C Web Annotation text-quote selector, is built to degrade gracefully: when it is not sure, it makes its best fuzzy guess.
For editing a regulatory submission, a best guess is the wrong default, because a misplaced edit is data corruption a reviewer might never catch. So we took that model and inverted its failure policy. The anchor widens its surrounding context in stages until the match is provably unique, and if it still cannot be certain, it stops. Verbatim from the source:
"We are EDITING a regulatory submission, where a misplaced edit is data corruption a reviewer may never catch. … the intended flag is edited, or nothing is — never a different one. Refusal is a reported outcome, never a silent no-op."
When that happens the writer does not get an error code. They get plain English: "This section has several identical flags here and the text has moved, so it isn't certain which one you picked. Reload the section and try again." Whatever you had typed is kept, not discarded. Most software treats "I'm not sure which one you meant" as a failure state to engineer out of existence. In regulated authoring it is often the correct answer, and the product should be willing to say it out loud.
One more rule is worth calling out, because it is the kind of thing that only gets written by someone who has watched a metric get gamed. Every flag is a tracked record that moves through open, resolved, dismissed, or stale, and two events deliberately never count as resolution. A flag that vanishes during a routine automatic save leaves the issue open until a writer-attributable sync actually sees the document. And regenerating a whole section marks its flags stale, not resolved, on a simple principle: a full rewrite is no evidence that anyone addressed the gap. Any AI writing tool could drive its gap count to zero by regenerating the section. We do not let that count, because the number is supposed to mean something. Every resolution also records how it happened — the writer's own edit, an agent edit via chat, or a bulk resolve — with the actor, the conversation, the transaction id of the document write, the section version at the time, and the timestamp, so those approve and reject decisions become part of the run bundle you can hand to QA.
Saying the true thing, even when it is smaller
There is a small detail in the gap card that captures the posture better than any feature does. The card under-claims, on purpose. Its text reads:
"This information wasn't found in the source passages Asthra drafted from — it may still exist elsewhere in the library. Ask Asthra to search further, replace it with your own text, or upload the source."
The flag's own registry description says the field is missing from the library. But that is more than the drafting agent actually knows: it searched the passages it retrieved, not every document in the library. So the card says the smaller, true thing, and points the writer at a deeper search that might still turn the value up. It is the same instinct behind why a confident, wrong metric is worse than no metric.
The honest version of "done"
None of this makes Asthra an AI that never gets stuck, and we are not trying to build that one. The drafting agent tells you where it got stuck, the resolution flow lets you point it at the answer, and Word shows you every word it changed before any of it counts. Flagging the gaps was the straightforward part. Filling them without asking you to take anything on trust is what took the work.
For the other end of this story, where the flags come from and how the FDA-grounded reviewer plants them in the first place, start with Regulatory Review: an FDA audit inside the draft. And because every resolution is recorded against the section's own record rather than the Word file — who filled the gap, from which source, in which document write — a flag's whole history stays auditable long after the edit.