Teaching an LLM where the buttons are
A crawler maps a web app so that AI-generated tests stop inventing selectors. Invented selectors went from 18 % to 0 %.
The problem
Ask a language model to write end-to-end tests for a web app and it will happily click on buttons that do not exist. It guesses selectors from what a page like this usually looks like, and the tests fail for reasons that have nothing to do with the app.
What I built
A crawler that walks through the app read-only, with every request that could change data blocked, and turns what it finds into a graph of states, transitions and selectors that were verified on the real page. The model only gets that graph as context when it writes tests. A dashboard shows the map, and a benchmark checks every generated selector against it.
What came out
With the graph, 0 % of the selectors were invented. Without it, 18 %. An independent judge model rated the generated tests above 0.9 for faithfulness with the graph and around 0.1 without.
Built with
Python, Playwright, Crawlee, FastAPI, LLMs via Ollama Cloud
Built at work for a client project, so no names or screenshots here.