Skip to main content

Command Palette

Search for a command to run...

Co-draft

Updated
•5 min read•View as Markdown
B
I build and study applied AI systems — LLM and RAG pipelines, agentic workflows, and the evaluation methods that tell us whether they actually work. I am a PhD researcher in Computer Engineering at Bahcesehir University, working on retrieval-augmented generation — standard, agentic, and graph-based — with a focus on evaluation, hallucination detection, and responsible AI. Alongside the PhD I have three first-author papers in progress, targeting EMNLP, ACL, and Information and Software Technology, on agentic-AI verification, responsible release of harmful-speech data, and the gap between "executable" and "correct" in solver-backed reasoning. In parallel, I design applied-AI prototypes at Turkish Airlines, turning LLM, ASR, and automation research into working proof-of-concept tools for high-stakes language assessment — including automated CEFR scoring pipelines benchmarked against human graders to catch model hallucination. My background is unusual on purpose. Alongside a First Class Honours BSc in Computer Science (Goldsmiths, University of London), I hold degrees in educational technology and language teaching and years of experience designing high-stakes assessments. That gives me a rare combination: I can build the model, evaluate it rigorously, and understand the human and pedagogical context it runs in. Interests: agentic RAG, LLM evaluation, hallucination detection, responsible AI, AI for education and assessment.

co-draft: let the model propose the edit, but keep the human holding the pen

The default way we use an LLM to edit text is "here's my document, rewrite it." You paste it back, it comes back changed, and you have no idea what actually moved. For a quick throwaway that's fine. For anything you're responsible for — a lesson, a spec, a public post — it's the wrong shape. The model is good at proposing edits and bad at knowing which edits you actually want, and "rewrite the whole thing" hides exactly the information you need to make that call.

co-draft is a small prototype built around the opposite default. You select the span you want changed and give an instruction. The model proposes a change to that span only. The change is shown as a before/after diff. Nothing is applied until you accept it; reject is a no-op; accepted edits go into a per-section history you can undo. The human stays in the loop by construction, not by discipline.

co-draft-diff

How it's put together

Four small pieces:

  • drafting turns raw text into titled sections (honouring Markdown headings, or grouping paragraphs when there aren't any).
  • refine is the loop: propose → review → apply. A proposal is just data until you call accept().
  • diff_engine does a word-level diff and renders a self-contained HTML artifact — two panes, highlighted add/remove/modify spans, and change statistics. No external CSS or JS; it opens straight from disk and adapts to a light or dark system theme.
  • providers is the model behind it, and there are three: a deterministic offline stub, an Ollama backend, and an OpenAI-compatible one. They share a two-method interface.

That stub matters more than it sounds. It's the default, it needs no GPU and no network, and it makes the whole pipeline reproducible — the committed sample artifacts regenerate byte-for-byte. It doesn't understand language; it maps instruction intents (concise, formal, simplify) to fixed transforms. It's a test double, so the repo runs anywhere in a couple of seconds, and you swap in a real model when you want real edits.

Here's a real one, from a local Gemma model rewriting a deliberately wordy lesson:

Before: Once the water reaches the ground it does a bunch of different things. After: Upon reaching the surface, the water undergoes several distinct processes.

The bug that only showed up with a real model

While generating that example I hit a good one. The stub's short edits looked fine, but the first real-model diff reported 9.3% similarity for an edit that was visibly about 93% the same text. The screenshot would have shipped a nonsense number.

The cause: the change statistics used Python's difflib.SequenceMatcher with its default autojunk=True. That heuristic treats characters appearing very often as "junk" to speed up matching on large inputs — and it quietly mangles the similarity ratio once a string gets past ~200 characters. The word-level diff itself already disabled autojunk; the statistics call didn't. One inconsistent default, invisible until the inputs got long enough, which only happened with real-model output.

The fix was one argument. The lesson is the older one: a default tuned for speed silently changed a correctness number, and it took real, messy input to expose it. Deterministic stubs are great for reproducibility, but they won't find the bugs that only live at the sizes real usage produces.

The point of the diff

The visual diff isn't decoration. LLMs introduce subtle errors even when you only ask them to "tighten" a sentence — a changed meaning, a dropped qualifier, an invented specific. If the model rewrites your whole document, those slip through. If every change is a small, highlighted, human-approved diff, you have a real chance to catch them. The design assumes the model is a fast but unreliable proposer, and puts the human where the judgment has to be.

That assumption is load-bearing. If you wire this into a pipeline that auto-accepts every proposal, you've removed the one thing it was built to protect.

What it isn't

  • Not production software: no auth, no persistence beyond files, single user, English-tuned tokenisation.
  • The diff is word-level, so very large rewrites render as big modified blocks rather than fine-grained changes.
  • The stub is a test double, not a writing assistant. Real edits need a real backend.

Why I built it

I work on LLM evaluation and human-AI teaming, and one thread of my research is about who verifies what when a model does open-ended work — where the human oversight actually has to sit, and why offline checks aren't enough. co-draft is a small, concrete take on that: a tool whose whole design is "the model proposes, the human decides, and the interface makes deciding possible."

Code: github.com/batinium/co-draft


Batın Örene — PhD researcher in AI (LLM evaluation, responsible AI, human-AI teaming). Papers under review at EMNLP and ACL. GitHub

6 views