Skip to content
Back to blog
engineering ai agents typst

How Our Template Agent Fixes Its Own Compile Errors

pdfs.build Team

Every AI code-generation demo shows the happy path: you ask for a thing, the model writes the thing, the thing works. The real product question starts one step later. What happens when the model’s output is almost right?

In pdfs.build, a template is three artifacts that must agree with each other: Typst code that lays out the document, a JSON schema that declares what data it accepts, and sample data that satisfies the schema. When you ask the chat to “add a due date under the invoice number”, an LLM edits all three. And a language model’s edit is plausible by construction, not valid by construction. A document API cannot return “almost compiles”.

Our answer is a rule we apply everywhere: the model never has the last word. The compiler does.

Post-passes: what runs after the model stops

After any turn that touches the document, a fixed sequence of passes runs before the result is considered done. Each pass has the same shape:

  1. A deterministic trigger decides whether the pass needs to do anything. No model call, no judgment, just a check.
  2. If triggered, a bounded repair runs. Sometimes that is mechanical rewriting; sometimes it is a narrowly scoped request back to the model. Bounded means a hard attempt limit, currently two.
  3. A verification step confirms the repair actually worked. The verifier is never the model. It is the compiler, or a concrete predicate on the document.

Three passes run today, in order: an image-reference pass (did the model declare an image field and then forget to place it?), a visual quality pass, and the compile-fix pass. The last one is the interesting one.

The compile-fix pass

If the turn modified the document, we compile it against the sample data. If that succeeds, we are done, and the freshly rendered first page becomes the template’s thumbnail as a side effect.

If it fails, we do not go to the model first. We apply a set of deterministic lint fixes: mechanical rewrites for the mistakes language models actually make in Typst, applied in one pass with no inference call. Then we compile again. If the mechanical fixes resolved the diagnostics, or even just reduced them, we keep that version and move on. A surprising share of failures ends here, at the cost of roughly zero.

Only when diagnostics survive the lint pass does the model get involved. We hand it the compiler’s diagnostics enriched with the surrounding source lines, so it sees the actual neighborhood of the error rather than a bare line number, and we ask for a repair. Then we compile again, because a repair claim from a model is worth exactly nothing until the compiler agrees.

Two attempts. If the document still does not compile after that, we stop and say so, with the diagnostics attached. An unbounded self-repair loop is a machine for converting money into increasingly creative versions of the same mistake.

Polish that can be rolled back

The visual quality pass has a different failure mode worth designing for. Its job is subjective: heuristics flag layout issues, or the user asked for polish, and the model gets one more shot at making the document look better without changing what it says.

“Better looking” and “still compiles” are independent properties. So the quality pass holds the last valid version of the document the whole time, and if a polish attempt introduces a compile error, it rolls back and tells the user exactly that: the polish was attempted, it broke something, the last valid version was preserved. The user never inherits a broken document from a pass whose entire purpose was cosmetic.

Why this shape

You could frame the whole system as a trust gradient. Deterministic checks are free and always right about what they check, so they run first and gate everything. Mechanical fixes are cheap and predictable, so they get the first shot at repairs. The model is powerful and unreliable, so it is invoked narrowly, fed precise context, capped at two attempts, and never trusted to grade its own work.

None of this is specific to PDFs. It is what “agent” has to mean for the output to be shippable: a model wrapped in a substrate that can verify its claims. The compiler happens to be a very good substrate. It is fast, it is deterministic, and it does not care how confident the model sounded.

If you want to see the loop from the outside, open any design in the template gallery, ask the chat for a change, and watch the passes run. The API reference covers what happens after the template is done: POST your JSON, get the PDF. And if you write Typst yourself, you can bring your own templates and skip the agent entirely.

Back to blog