Co-draft
co-draft: let the model propose the edit, but keep the human holding the pen
The default way we use an LLM to edit text is "here's my document, rewrite it." You paste it back, it comes back changed, and you have no idea what actually moved. For a quick throwaway that's fine. For anything you're responsible for — a lesson, a spec, a public post — it's the wrong shape. The model is good at proposing edits and bad at knowing which edits you actually want, and "rewrite the whole thing" hides exactly the information you need to make that call.
co-draft is a small prototype built around the opposite default. You select the span you want changed and give an instruction. The model proposes a change to that span only. The change is shown as a before/after diff. Nothing is applied until you accept it; reject is a no-op; accepted edits go into a per-section history you can undo. The human stays in the loop by construction, not by discipline.

How it's put together
Four small pieces:
- drafting turns raw text into titled sections (honouring Markdown headings, or grouping paragraphs when there aren't any).
- refine is the loop: propose → review → apply. A proposal is just data until you call
accept(). - diff_engine does a word-level diff and renders a self-contained HTML artifact — two panes, highlighted add/remove/modify spans, and change statistics. No external CSS or JS; it opens straight from disk and adapts to a light or dark system theme.
- providers is the model behind it, and there are three: a deterministic offline stub, an Ollama backend, and an OpenAI-compatible one. They share a two-method interface.
That stub matters more than it sounds. It's the default, it needs no GPU and no network, and it makes the whole pipeline reproducible — the committed sample artifacts regenerate byte-for-byte. It doesn't understand language; it maps instruction intents (concise, formal, simplify) to fixed transforms. It's a test double, so the repo runs anywhere in a couple of seconds, and you swap in a real model when you want real edits.
Here's a real one, from a local Gemma model rewriting a deliberately wordy lesson:
Before: Once the water reaches the ground it does a bunch of different things. After: Upon reaching the surface, the water undergoes several distinct processes.
The bug that only showed up with a real model
While generating that example I hit a good one. The stub's short edits looked fine, but the first real-model diff reported 9.3% similarity for an edit that was visibly about 93% the same text. The screenshot would have shipped a nonsense number.
The cause: the change statistics used Python's difflib.SequenceMatcher with its default autojunk=True. That heuristic treats characters appearing very often as "junk" to speed up matching on large inputs — and it quietly mangles the similarity ratio once a string gets past ~200 characters. The word-level diff itself already disabled autojunk; the statistics call didn't. One inconsistent default, invisible until the inputs got long enough, which only happened with real-model output.
The fix was one argument. The lesson is the older one: a default tuned for speed silently changed a correctness number, and it took real, messy input to expose it. Deterministic stubs are great for reproducibility, but they won't find the bugs that only live at the sizes real usage produces.
The point of the diff
The visual diff isn't decoration. LLMs introduce subtle errors even when you only ask them to "tighten" a sentence — a changed meaning, a dropped qualifier, an invented specific. If the model rewrites your whole document, those slip through. If every change is a small, highlighted, human-approved diff, you have a real chance to catch them. The design assumes the model is a fast but unreliable proposer, and puts the human where the judgment has to be.
That assumption is load-bearing. If you wire this into a pipeline that auto-accepts every proposal, you've removed the one thing it was built to protect.
What it isn't
- Not production software: no auth, no persistence beyond files, single user, English-tuned tokenisation.
- The diff is word-level, so very large rewrites render as big modified blocks rather than fine-grained changes.
- The stub is a test double, not a writing assistant. Real edits need a real backend.
Why I built it
I work on LLM evaluation and human-AI teaming, and one thread of my research is about who verifies what when a model does open-ended work — where the human oversight actually has to sit, and why offline checks aren't enough. co-draft is a small, concrete take on that: a tool whose whole design is "the model proposes, the human decides, and the interface makes deciding possible."
Code: github.com/batinium/co-draft
Batın Örene — PhD researcher in AI (LLM evaluation, responsible AI, human-AI teaming). Papers under review at EMNLP and ACL. GitHub
