A Python UI framework that AI models learn from the README alone
ui is a tiny Streamlit-style framework for personal tools. You write a page top to bottom in Python, every ui.* call renders HTML in place, and every interaction re-runs the whole page while htmx swaps only what changed. State is a file, cache is a file, and every route is a file. It’s standard library only, with no build step and no node_modules.
import ui
"# Tip splitter"
bill = ui.number("bill", value=40)
people = ui.slider("people", 1, 12, value=2)
ui.stat(f"${bill / people:.2f}", "each")
That’s a whole app. There are no callbacks to wire and no state to declare. Drag the slider and the number updates.
Letting a small model find the gaps
The README is written for AI models as much as for people, so it has to be good enough for a model to write correct code from it with nothing else. I tested that directly:
- Give Claude Haiku the README as its only system prompt, with zero tools and an empty working folder, so it can’t read the framework’s source even by accident.
- Ask it to build an app from a short brief.
- Grade the app in a real browser with flow tests written from the brief’s requirements, not from what the code happens to do. A smoke test passes while a Delete button silently does nothing.
- Ask “why” five times about each failure, then make one minimal fix to the README or the engine.
- Repeat until three good apps in a row.
It ran twice, first with Gemma for thirteen rounds, then with Haiku, which converged on rounds eight through ten. The failures come in two layers. First come gaps in the docs, where the model guesses an unwritten rule and loses. Once the docs are good, engine bugs show up, where the model writes code that’s correct by the docs and the framework breaks it.
Tournaments
To find out what the framework is good at, I ran elimination tournaments. For a theme like computer vision, geography maps, small LLM apps or one-player games, eight agents each build an app, a three-judge panel scores them, and the field narrows from eight to four to two to one. Survivors rebuild each round with the judges’ critique in hand. Besides quality, the judges scored how well each entry used the framework’s built-ins, which shows which primitives carry their weight.
This version exists to find out whether the syntax feels right. A Rust engine comes later, with the same page code.