Featured project: BenchFlow (Open Source)
Contributed to RL runtime environments and community-based AI benchmarks.
Python · LLM evaluation · AI agents
A collection of projects I've worked on, ranging from academic research to personal explorations. If you think it would be useful or interesting to collaborate on a project, please contact me to discuss.
Try clearing the skill filter.
Contributed to RL runtime environments and community-based AI benchmarks.
Python · LLM evaluation · AI agents
Analysis of LLM factuality, hallucination, topic patterns, and cross-cultural reasoning failures on multicultural riddles.
LLM evaluation · Leadership · Multicultural AI
Studying how monitor effectiveness changes as the capability gap between monitor and target models widens, with a case study on distinguishing sandbagging from genuine incapability.
Python · LLM evaluation · Chain-of-thought monitoring
Created and maintain a personal (or group) academic Astro theme for research portfolios, publications, projects, and blogs.
Web design · TypeScript · Astro
Extended Tau2-Bench for low-resource Southeast Asian languages and localized agentic evaluation.
Python · LLM evaluation · AI agents · Multilingual AI
Minimal Llama-2 implementation in PyTorch with RoPE, RMSNorm, SwiGLU, self-attention, and transformer blocks.
Python · PyTorch · Model training
Designed the SEACrowd website and managed social content for a Southeast Asian AI research community.
Web design · JavaScript · Jekyll · Bootstrap
Evaluated whether implicit values in agentic LLMs are inclusive of animals, by extending CaML's The Animal Compassion
LLM evaluation · AI agents · AI x animals
Investigated multilingual LLM representations with SAEs and feature steering across 67 languages.
Python · PyTorch · Mechanistic interpretability
Luma-like platform for discovering, managing, and RSVPing to local recreational sport events.
TypeScript · React · PostgreSQL
AI chatbot generating memorable English and Mandarin mnemonics with QLoRA fine-tuning and DPO preference modeling.
Python · Fine-tuning · PyTorch
Question-answering system over Obsidian-style personal notes using LlamaIndex, OpenAI API, and retrieval-augmented generation.
Python · RAG · LlamaIndex
Replicated and extended a synthetic control analysis of Philadelphia SNAP benefit redemption in R.
R · Causal inference · Synthetic control
Analyze GP visit patterns using Zero-Inflated Poisson models with complete and partial pooling, with Bayesian inference and data imputation.
Python · PyMC · Bayesian modeling · Data imputation
Text classification system to automatically classify notes and assignments using SVM and Naive Bayes with 87% accuracy.
Python · scikit-learn
Languages, frameworks, and infrastructure from my work.
Loading the latest project details from GitHub…
Recent public activity from @mychiffonn . Hover over a day, or focus the heatmap and use the arrow keys, to inspect the past year.
Syncing with GitHub…