Using LLMs in a publishing pipeline where they help, and refusing them where they hallucinate. Scoped as a standalone workstream or folded into a full data scientists build.
from $1,500 as a standalone workstream · included in the Data Scientists Build tier
Language models are excellent at normalising and summarising known facts and dangerous at supplying unknown ones. We run this as a fixed-scope workstream: 4 defined deliverables, one named approver, and a written handover at the end. It attaches to an existing data scientists engagement or stands alone if that is the only piece you are missing.
Permitted uses: normalisation, summarisation, classification, extraction
Forbidden uses: fact generation, statistic invention, source fabrication
Human-verified sampling rates per task type
Evaluation harness for output quality drift over time
We look at what data scientists already hold — systems, exports, APIs — and score each axis for demand and defensibility.
The data contract is written and the first template is designed against real rows, not placeholders.
4,500–10,500 URLs published with schema, internal links, sitemap entries and IndexNow.
Indexation and impression data decides what widens and what gets cut. Templates, gates and runbook transfer to you.
Start at 5% per template family per batch, and raise it whenever the error rate exceeds your accepted threshold.
We'll audit the data source, size the first batch, set the performance budget and tell you honestly if programmatic is the wrong tool for your category.