Projects

A collection of projects I've worked on, ranging from academic research to personal explorations. If you think it would be useful or interesting to collaborate on a project, please contact me to discuss.

Skill
Show 23 more skills Hide additional skills

Studying how monitor effectiveness changes as the capability gap between monitor and target models widens, with a case study on distinguishing sandbagging from genuine incapability.

Python · LLM evaluation · Chain-of-thought monitoring

Extended Tau2-Bench for low-resource Southeast Asian languages and localized agentic evaluation.

Python · LLM evaluation · AI agents · Multilingual AI

Tech Stack

Languages, frameworks, and infrastructure from my work.

Languages

  • Python
  • TypeScript
  • JavaScript
  • SQL
  • Bash
  • R

AI & ML

  • LLM
  • prompt engineering
  • evaluation
  • agents
  • RAG
  • Fine-tuning
  • PyTorch
  • Hugging Face ecosystem
  • Unsloth
  • Inspect AI

Backend & Data

  • FastAPI
  • Flask
  • Express.js
  • REST APIs
  • PostgreSQL
  • SQLite
  • ClickHouse
  • Prisma

Frontend & Product

  • React
  • Astro
  • Next.js
  • Tailwind CSS
  • shadcn/ui

DevOps

  • Git
  • Docker
  • CI/CD
  • Vercel
  • Railway

On GitHub

mychiffonn / website

merged PRs by @mychiffonn

Loading the latest project details from GitHub…

stars forks open issues and pull requests
Syncing with GitHub…

GitHub Activity

Recent public activity from @mychiffonn . Hover over a day, or focus the heatmap and use the arrow keys, to inspect the past year.

Syncing with GitHub…