GeopolAI - measuring bias across GPT, Claude and Gemini
The challenge
LLMs are used to summarize serious conflicts, but few teams measure whether different models recommend different levels of escalation, civilian-risk framing or realpolitik.

What I built
GeopolAI runs the same geopolitical scenarios through several LLMs and scores their divergence. It is a local R&D demonstrator, not a deployed client product.
It is useful because it makes model disagreement visible. The goal is not to decide which model is right, but to stop pretending they are interchangeable.
Key engineering points
Parallel scenario submission to GPT, Claude and Gemini.
Scoring axes for escalation, civilian risk and realpolitik framing.
Comparison UI for seeing divergence instead of reading isolated answers.
Explicit positioning as audit R&D, not an automated decision system.
Similar technical risk?
I can help scope the risk, architecture and first deliverable.
A 30-minute first call is enough to see whether I am the right profile for the problem.
A similar challenge?
A system like this one to build? Let's talk.
I take on critical technical work — from scoping to production, no debt or lock-in once it's handed over. Fastest way to see if it fits: a 30-minute call.
I reply within 24h — often sooner.