Skip to content
Philippos MelikidisStuttgart region, October 2026

2026 · client work

Teaching an LLM where the buttons are

A crawler maps a web app so that AI-generated tests stop inventing selectors. Invented selectors went from 18 % to 0 %.

The problem

Ask a language model to write end-to-end tests for a web app and it will happily click on buttons that do not exist. It guesses selectors from what a page like this usually looks like, and the tests fail for reasons that have nothing to do with the app.

What I built

A crawler that walks through the app read-only, with every request that could change data blocked, and turns what it finds into a graph of states, transitions and selectors that were verified on the real page. The model only gets that graph as context when it writes tests. A dashboard shows the map, and a benchmark checks every generated selector against it.

What came out

With the graph, 0 % of the selectors were invented. Without it, 18 %. An independent judge model rated the generated tests above 0.9 for faithfulness with the graph and around 0.1 without.

Built with

Python, Playwright, Crawlee, FastAPI, LLMs via Ollama Cloud

Built at work for a client project, so no names or screenshots here.

← All work