GPS-Bench: AI Governance Policy Simulation

Featured
to Present Research Open Source
AI governance LLM evaluation Multi-agent systems

I co-author GPS-Bench, an evidence-grounded benchmark and simulation framework for evaluating AI-governance policy forecasts against historical records.

The framework represents policies as heterogeneous temporal graphs connecting antecedent events, policy provisions, actor states and actions, world-state changes, and stakeholder impacts. It evaluates legislative passage, affected-actor identification, actor-action prediction, and actor-level impact direction under temporal evaluation.