I co-author GPS-Bench, an evidence-grounded benchmark and simulation framework for evaluating AI-governance policy forecasts against historical records.
The framework represents policies as heterogeneous temporal graphs connecting antecedent events, policy provisions, actor states and actions, world-state changes, and stakeholder impacts. It evaluates legislative passage, affected-actor identification, actor-action prediction, and actor-level impact direction under temporal evaluation.