Springer Best Paper Award · ISVC 2022
Learning When to Say “I Don’t Know”
Per-class abstention thresholds that need no rejection cost or coverage target: CIFAR-100 selective accuracy climbs from 88.3% to 97.8% at 77.3% coverage.
01Machine learning · Research & engineering
PhD, Ohio State (2026). I built LoRA-trained policies that decide when a RAG system should revise its answer, evaluated on 25,870 questions, and a reject-option method that won a Springer Best Paper Award.
FIG.0 1,000 CIFAR-100 test predictions, grown to their confidence. Below the line, the model says “I don’t know.” Try the threshold
Available nowfor Research Scientist,Research Engineer,ML Engineer,Applied Scientist,SWE (AI)
SF Bay Area / NYC.U.S. citizen
LoRA / PEFT Teacher-model synthetic data vLLM HF Transformers
LoRA-trained scorers that choose answer, revise, or abstain; evidence-grounded QA models trained on teacher-generated data at DCS. Evidence: RAG revision DCS role
DPR dense retrieval BM25 MonoT5 reranking FAISS
Three Wikipedia retrieval setups behind a 25,870-question evaluation. Evidence: RAG revision
Paired outcomes Seeds + paired t-tests Bootstrap intervals LLM-as-judge Selective prediction Calibration (ECE, Brier)
3 seeds with paired t-tests and 10,000-replicate bootstraps; a reimplementation checked against its reference by 37 parity tests. Evidence: RAG revision Reject option
ViTs Linear probes on DINOv2, DINOv3, SigLIP 2 Self-supervised pretraining
Probes on four frozen ViT backbones; self-supervised vision models at AFRL in 2024. Evidence: Backbone re-run AFRL 2024
Python PyTorch NumPy Slurm Singularity uv Git Linux HPC TypeScript
Ported the B-CDF rule from Python to TypeScript for the live demo; Slurm GPU jobs for the RAG study. Evidence: Live demo RAG revision
2.1 Selected work
Dissertation capstone
Revision can repair a wrong answer or harm a right one. The system has to choose before paying for the full revision.
What I built
Unpublished Unpublished. Revised after ACL Rolling Review; arXiv version in preparation. Code and artifacts will not be released.
ISVC 2022 · MVA 2025 · 2026
A classifier should decline its least reliable predictions without a hand-set rejection cost or coverage target.
The method · 2022
The 2026 re-run
The demo
Best Paper · Code public Method peer-reviewed: Springer Best Paper Award at ISVC 2022, journal extension in Machine Vision and Applications (2025). The 2026 re-run is public code, not peer-reviewed.
WMT 2024
Translation scores alone can’t show whether a model reads the image or ignores it.
What I built
Peer-reviewed Peer-reviewed, WMT 2024.
Also public
ICCV 2021 Workshop · Code public
An OpenStreetMap-guided Python pipeline that extracts candidate construction-site regions and downloads imagery over time. Rebuilt in 2026 as one package with swappable history and imagery backends and a 410-test suite.
Code public
PyTorch utilities for histogram binning and global and class-wise temperature scaling, with expected-calibration-error summaries and calibration plots.
2.2Live demo · ISVC 2022 Best Paper
Sort a classifier’s guesses by confidence and decline the least sure ones. The rule from Learning When to Say “I Don’t Know” keeps declining only while the declined guesses are no better than a coin flip. The paper fits a line per class; this demo uses one line for all 100 classes. Drag it.
Guess
Truth
·
1,000 of the 10,000 test images, one dot each, placed by the model’s confidence. Mistakes stack at the bottom of each column; hollow dots are declined. The dashed line is where the rule put the threshold. Hover or press the arrow keys to inspect an image.
No better than a coin flip: the declined guesses were right 50.6% of the time, so declining them gives up almost nothing.
Model: DINOv2 ViT-S/14 (frozen) + linear probe, temperature-scaled. Data: CIFAR-100. The line is learned on 10,000 held-out training images; every number here is measured on the 10,000 test images, of which the plot shows 1,000.
03 Research
Springer Best Paper Award · ISVC 2022
Per-class abstention thresholds that need no rejection cost or coverage target: CIFAR-100 selective accuracy climbs from 88.3% to 97.8% at 77.3% coverage.
04 Experience
05 Get in touch
Research Scientist, Research Engineer, ML Engineer, Applied Scientist, and SWE (AI) roles in the San Francisco Bay Area or New York City.