Newsletter conversion: project update
4 September 2026

Newsletter PDF-to-HTML conversion

A short account of what has been completed, why the design work took longer than expected, and the approach we are testing now.

Current position

The content foundation exists. The design method is still being proved.

Across all 63 newsletters, we have extracted text, images, URLs and much of the relationship between them. Text and URLs are the most reliable parts. Image crops and some image-to-link associations still need visual review.

What remains

We need a repeatable way to turn this material into a good responsive page. A generated page or a technically valid page is not automatically ready to publish.

01

What we tried

First approach · about one month

A fixed publishing format

We first placed the original story panels at the top of the page and the searchable text below them. This made the source material visible and the text available on the web. To support this approach, we also built an internal tool for creating, checking and correcting these pages.

What gave us confidence

The May 31 version from the client repository made the original panels easy to browse, with a readable HTML edition below.

How the project changed

After we showed this version, the project took a different path: using an LLM to assemble the extracted text, images and links into a designed page closer to the PDF. We therefore moved away from the fixed panel-and-text format.

Second approach

LLM design constrained by the PDF

We then asked models to reproduce the PDF more closely. To help them, we added coordinates, crop choices, HTML fragments and layout instructions to the manifest. This carried the old design into the input and left the model trying to redesign an already prescribed layout. The content was often present, but spacing, typography, crops and link placement remained inconsistent.

What gave us confidence

After careful tuning, the June result came close enough to suggest that reconstruction might work.

Where it broke down

The method did not transfer reliably across models. Later attempts still had inconsistent spacing, typography, crops and text placement.

Why checking did not solve it

The automated approval tool did not work reliably

Every version still needed manual checking across stories, links, languages and screen sizes. We built an automated visual checker, but it could recognise that two pages looked different more reliably than it could decide whether the difference was an actual error. It raised too many false alarms to approve newsletters on its own.

02

Current test: verified content with less design instruction

We now keep the content, hierarchy, order and ownership, but remove PDF coordinates, old HTML and renderer choices. The model is free to design the page.

Test case: December 31 English, one of the more difficult newsletters. The same lean source was used for the three outputs below.

Finding: all three fit within a 390 px mobile viewport. This is an improvement over the earlier experiment. It proves containment, not content accuracy, crop quality, link ownership or publication readiness.

The Fable 5.1 output is the most promising design we have seen. This suggests that a lean, verified content package gives the model a better starting point than a layout-heavy manifest. We still need to establish whether the result can be reproduced across different newsletters and what each conversion will cost.

03

Proposed next step

First reproduce the Fable 5.1 result on a small, varied English/Hindi sample and record the real cost per edition. Its first result was materially more accurate, so we expect the remaining design work to be significantly less than the earlier template and reconstruction work. The sample will tell us whether that improvement holds across editions.

The decision needed is whether the goal is a readable, accurate web edition that may differ from the PDF, or close reproduction of each PDF. The first now looks practical. The second has not worked reliably in our tests.

Tentative timeline

  1. By 8 September: run Fable 5.1 on four representative English/Hindi editions and record quality, repeatability and cost.
  2. By 11 September: apply only the recurring fixes found in that sample, rerun it, and share a go-or-stop recommendation.

At that point, we will have enough evidence to propose a realistic schedule for the remaining newsletters. These dates assume continued access to Fable 5.1 and timely human review.