A daily public log of AI model drift: the same 50 frozen questions hit GPT, Claude and Gemini, answers kept word-for-word and compared day by day.
modeldrift.watch is a daily public log of how AI models shift their answers. Since September 10, the same 50 frozen questions are sent to GPT, Claude, and Gemini through their APIs every day. Every answer is saved exactly as typed, scored with code, and compared against prior days. If a model stops agreeing on a fact, skips a question, or changes a suggestion, it shows up with the full before and after. The archive is append-only — no backfilling. As of today it carries 3,311 answers on record and was updated this morning.
The exhibit format is what makes it land. One entry tracks a legal question over two days: on September 10 Gemini correctly answered that a cited court case "is not a real case"; on September 11 the same model cited it as binding precedent. Same question, 11 hours apart, both answers kept word for word. It's a quietly alarming demonstration of why "the model said so" is not a citation.
The author built it solo and writes plainly: "built solo, claude did most of the coding." A one-person observation instrument for the era of silent model updates — part accountability journalism, part science project, all public.