prompt-contracts
Contract tests for LLM prompts, so a model update cannot quietly break an integration.
The problem
When a provider updates a model, changes a default or when you switch to a local model, prompts that worked yesterday can return something slightly different today. Nothing fails loudly, the integration just gets worse.
What I built
A specification and toolkit for writing contracts against prompt outputs, covering structure, stability and consistency, and running them like tests. Comparisons come with proper statistics, Wilson and Jeffreys intervals and McNemar tests, so a difference is only called a difference when it is one.
What came out
Released on PyPI. It became the starting point for how I evaluate LLM systems in general, including the thesis.
Built with
Python, published on PyPI