Projects

First author · under review at ICLR 2027 Python · PyTorch · Transformers · scikit-learn

Audit-Prune: removing shortcuts from LLM benchmarks

Six answer-only surface features score 68.9% on binary-choice TruthfulQA, more than 16 of the 18 models on the llm-stats leaderboard, and several of 13 other benchmarks leak too. Audit-Prune cuts TruthfulQA's leak to near chance while keeping model rankings (Spearman 0.915).

Audit-Prune: a probe that judges answers by their cover Five pairs of closed books stand for true and false answer pairs from TruthfulQA; the black bars above them are the questions, which the probe never reads. A magnifier, the probe, flags the two pairs in which one cover opens with the word No. Audit-Prune drops those pairs, and on its second look the probe is left with a question mark. The readouts give the reported results: surface-audit AUC falls from 0.715 to 0.528 (chance is 0.5), 476 of 790 pairs are kept, and model ranking is preserved with Spearman rho 0.915. The books and their covers are a schematic illustration, not measured data. Questions: never read No, No, ! ? Probe: opens with “No”? Probably the true one. Audit-Prune drops the pairs that leak most. Giveaway pairs pruned: the probe is near chance. Surface-audit AUC 0.715 0.528 chance 0.5 was 0.715 Pairs kept 790 476 of 790 Model ranking preserved: Spearman ρ = 0.915
Authors listed alphabetically · under review at FoDS C++ · pybind11 · Python

dKS: a fast test for whether two datasets differ

Extends the Kolmogorov–Smirnov distance to multiple dimensions as a unit-invariant metric. The approximation runs about 76,000× faster than exact computation on a million 2-D points; I wrote the open-source library.

dKS: the largest gap over unbounded rectangles Schematic, not measured data: two samples of points, P in blue and Q in orange, on a grid. A corner hops between grid corners; the unbounded rectangle of everything below and to the left of it is highlighted while two bars show the share of each sample inside, and it settles on the corner where the gap between the bars is largest, which is the dKS distance. Squeezing the horizontal axis, as when grams become kilograms, leaves both bars exactly unchanged. sample P sample Q share inside gapdKS PQ gramskilograms two samples: are they the same? at every grid corner, measure the gap largest gap over all corners = dKS grams → kilograms: exactly unchanged
Creator · used in 2 papers

Public LLM datasets on Hugging Face

Three contrast sets (3,053 pairs) for steering LLMs, each isolating one behavior (conciseness, positivity, formality), two built with Martian AI; plus TruthfulQA-476 and two held-out sets for testing truthfulness classifiers.

Contrast pairs on Hugging Face become a steering vector Schematic of a contrast dataset. One prompt has an informal (orange) and a formal (blue) answer with the same meaning; many such pairs form two clouds of dots, the arrow between the two cloud means is the steering vector, and adding it rolls the model's output along the axis from informal to formal with the weights untouched. The wording, the dots and the distance the output travels are an illustration, not measured data; the pair counts (994, 1,059 and 1,000) are the real sizes of the three sets, which are public on Hugging Face: a bracket labelled with the hugging face emoji and the words "on Hugging Face" groups them. informal formal How was the demo? nailed it, lol It went very well. same meaning steering vector difference of the two means nailed it, lol It went very well. weights untouched 🤗 on Hugging Face conciseness 994 pairs positivity 1,059 pairs formality 1,000 pairs
First author · EAI SmartSP 2025 (Springer) Flutter · BLE · Node.js · MongoDB

MotionPI: data platform for an NIH wearable study

Wristbands detect moderate-to-vigorous activity and trigger in-the-moment phone surveys (113 participants enrolled to date). As sole developer of the app, backend and offline-first sync, I built a pipeline that handles ~7.7M records a day with zero malformed writes.

MotionPI: wristbands to phone to database Two wristbands stream data packets over BLE to a phone, which passes them on to a MongoDB database. When the link to the server opens like a drawbridge, packets wait in a buffer on the phone and are sent once the link closes again. A wristband ENMO trace rises above the 100.6 mg line for 4.9 of 7 minutes, a trigger bit travels to the phone, and a survey card appears. The trace, packets and survey faces are a schematic illustration, not measured data; the figures in the labels are from the project write-up. L R Survey BLE trigger bit sync when safe offline 2 wristbands writes locally first MongoDB wristband ENMO 100.6 mg 4.9 of 7 min above ~7.7M records a day zero malformed writes
First author · SIGSPATIAL 2026, accepted (23% rate) Python · pyscan

Hotspot detection on aggregated health data

Fast hotspot detectors need points, not region totals. Sampling 20–50 points per region instead of one improves detection power and scans the contiguous U.S. in 0.33 s (FlexScan: 1,109 s with its significance test).

Centroids versus sampled points for hotspot detection A stylised map of twelve regions with a hidden hotspot where three regions meet. With one centroid dot per region the scanning lens finds no points there and misses it; once every region is replaced by 20 to 50 points sampled inside its outline (drawn heavier where a region has more cases to share), the lens locks onto the hotspot. The map, the hotspot and the points are a schematic illustration, not measured data. ? The shortcut: 1 centroid per region Instead: 20–50 sampled points per region centroidmore caseshotspot scan misses it scan finds it