Publications and manuscripts
Twelve studies on one question: whether machine learning results mean what they claim. 1 published, 5 under review, 2 working papers, 3 manuscripts and 1 in preparation. Every status is stated exactly and never upgraded.
Showing all 12
Benchmark Accuracy Is Not an Identified Quantity
Most orderings on a published MATH-Hard leaderboard cannot be separated once the responses a scorer could not read are counted.
- Benchmark Evaluation
- NLP · LLMs
- Deployment Shift
Higher AUC, Fewer Cases Flagged: Subgroup Fairness Metrics Can Reverse at the Decision Threshold in Federated Diabetes Screening
Federated training narrows the White-Black AUC gap while a race-blind screening policy flags fewer Black respondents.
- Fairness
- Healthcare AI
- Privacy · Federated
Benchmarks Do Not Report Whether They Could Read the Answer
Two answer extractors shipped in the same evaluation harness disagree by 88.5 points on the same 400 MMLU responses.
- Benchmark Evaluation
- NLP · LLMs
Scoring-Rule Sensitivity in LLM Evaluation: Evidence from Indian Financial Regulatory Text
Changing only the scoring rule reorders twelve LLMs on a new Indian financial-regulation benchmark, and no model keeps its rank.
- Benchmark Evaluation
- NLP · LLMs
- Financial ML
Comprehensive Evaluation of Machine Learning for Type 2 Diabetes Risk Prediction: Large-Scale External Validation and Fairness Analysis
Internally validated diabetes models lose discrimination on 1.28M external records, and adults over 60 lose the most.
- Healthcare AI
- Fairness
- Deployment Shift
TrustShift: A Mechanism-Aware Audit of Machine-Learning Deployment Shift
Across four real deployment shifts, shift magnitude alone does not consistently explain which trustworthiness axis fails.
- Deployment Shift
- Healthcare AI
- NLP · LLMs
Confidently Wrong: Ranking Inversion in Cross-Network Denial-of-Service Detection
On a new network, DoS detectors do not decay toward chance: many pass through it and score attacks below benign traffic.
- Network Security
- Deployment Shift
Scorer-Induced Partial Identification of Benchmark Comparisons: An Audit-Design Approach
Once a scorer’s error rate is audited rather than assumed, audit uncertainty outweighs sampling uncertainty in all 120 well-defined comparisons.
- Benchmark Evaluation
- NLP · LLMs
Persistent Racial Disparities in U.S. Mortgage Approval: Evidence from 42 Million Applications, 2020-2024
The public-data Black-White mortgage approval gap at national scale: how large, where it lives, and how it responds to institutional boundaries.
- Fairness
- Causal Inference
Who Bears the Burden? Heterogeneous Racial Approval Differentials in U.S. Mortgage Lending: Causal Forest DML on 42 Million HMDA Applications
Not whether an average penalty exists, but who bears it, and through which underwriting channel.
- Causal Inference
- Fairness
When the Gate Stays Closed: Empirical Evidence of Near-Zero Cross-Sectional Predictability in Large-Cap NASDAQ Equities Using an IC-Gated Machine Learning Framework
A deployment gate for financial ML, and the discipline to report that it stayed closed.
- Financial ML
- Deployment Shift
Text Genre, Not Platform Identity, Predicts Transfer Failure in Mental Health Natural Language Processing: A Five-Axis Deployment Audit Across Five Corpora
A five-axis pre-deployment audit of mental-health text classifiers moved across platforms and corpora.
- NLP · LLMs
- Fairness
- Healthcare AI