Skip to content
Philippos MelikidisStuttgart region, October 2026

2025 · open source

prompt-contracts

Contract tests for LLM prompts, so a model update cannot quietly break an integration.

The problem

When a provider updates a model, changes a default or when you switch to a local model, prompts that worked yesterday can return something slightly different today. Nothing fails loudly, the integration just gets worse.

What I built

A specification and toolkit for writing contracts against prompt outputs, covering structure, stability and consistency, and running them like tests. Comparisons come with proper statistics, Wilson and Jeffreys intervals and McNemar tests, so a difference is only called a difference when it is one.

What came out

Released on PyPI. It became the starting point for how I evaluate LLM systems in general, including the thesis.

Built with

Python, published on PyPI

Source on GitHub

← All work