Projects

Client and independent case studies organized around the decision, stakeholder constraint, data, method, result, and operating limit.

Start with the National Health Workforce Registry or Supporter Survey to Work Plan for client delivery. Independent systems follow with their baselines, outcomes, and limits.

All projects

Real client projects come first, followed by independent technical work with measured evidence and explicit limitations.

10 projects shown

National Health Workforce Registry

Client delivery Public-sector data

Governed registry, API, and public dashboard for a national health regulator, designed to expose provenance and refuse unsupported output.

Delivered: an official catalogue covering 8,627 facilities, 48 traced requirements, release certification, disclosure controls, and a fail-closed public interface.

1,132 high-confidence external OSM location matches remain separate from official identity records and await field verification; 7,495 facilities remain explicitly unmapped. Practitioner records are synthetic pending approved source data.

  • Data governance
  • Python API
  • Disclosure controls

Supporter Survey to Work Plan

Client analysis Confidential survey data

An anonymized volunteer project that turned open-ended supporter feedback into a practical work plan for a one-person climate advocacy nonprofit.

Delivered: seven reviewable themes, three recommendations tied to counts, and a follow-on accessibility and performance audit ordered by effort.

89 survey responses, including 26 substantive website comments. The coding is exploratory, no implementation impact was measured, and no verbatim response left the protected source workbook.

  • Survey analysis
  • Qualitative coding
  • Accessibility review

Dispatch Optimizer

Decision optimization Synthetic benchmark

Decision optimizer for a synthetic field-service operation, assigning technicians under skill, capacity, service-level, travel, and overtime constraints.

Result: greedy plus 2-opt beat naive on 99% of randomized scenarios and lost on none. Warm-started CP-SAT reached median parity with that stronger heuristic at the shipped eight-second budget.

The interface image shows one seeded synthetic day. Written evidence covers 200 randomized scenarios and a separate 24-scenario solver-budget sweep.

  • OR-Tools CP-SAT
  • FastAPI
  • Next.js
  • Postgres
  • Optimization
Optimizer Results interface for one seeded synthetic dispatch day, showing recommended technician routes and job timing

Water System Risk & Funding Priority Index

Risk modeling Ohio public data

Explainable Ohio public drinking water screening model using public data, transparent scoring, service-area boundaries, source-water overlays, API-backed search, and dashboard delivery.

Result: the weighted index scored 0.740 ROC AUC versus 0.734 for simply counting prior violations, while its compliance component alone reached 0.781.

16,339 Ohio public records. The composite did not earn its complexity; geography is approximate and one funding term remains incomplete.

  • FastAPI
  • Postgres
  • Python
  • GIS
  • Public data
Water System Risk dashboard with Ohio screening map and review tier charts

Automotive Analyst

Applied AI Synthetic manufacturing data

Bring-your-own-key text-to-SQL agent for the synthetic factory warehouse, with client-side model calls, schema grounding, read-only SQL guardrails, and visible query evidence.

Result: an adversarial suite exposed and drove a fix for a quoted-identifier bypass; the final run blocked 43 of 43 attacks, accepted 16 of 16 legitimate queries, and stored zero model keys on the backend.

Synthetic warehouse. The guardrail evaluation does not establish that model-generated SQL answers the user's question correctly.

  • Next.js
  • FastAPI
  • PostgreSQL
  • Text-to-SQL
  • AI safety
Automotive Analyst text-to-SQL app with BYOK panel, sample questions, and guardrail workflow

Black Box AI

Applied AI · Public data

Natural-language aviation safety analysis over NTSB final reports, combining guarded SQL, cited narrative retrieval, validated charts, visible queries, and audit trails.

Result: hybrid retrieval reached 0.902 MRR versus 0.853 for semantic search, but the exploratory evaluation did not justify the added complexity, so semantic search remains the default.

7,462 public NTSB reports with visible SQL and citations. The labeled evaluation contains only 17 questions.

  • Guarded SQL
  • Hybrid retrieval
  • Citations
  • FastAPI
Black Box AI chart and SQL-backed aviation safety answer

Manufacturing Intelligence Platform

Decision dashboard · Synthetic data

Factory analysis using a reproducible synthetic assembly dataset, PostgreSQL star schema, validation checks, and an executive decision dashboard.

Result: the pipeline recovered deliberately seeded process signals, including 69 percent downstream defect propagation and a 45 percent seeded defect reduction.

35.3 million synthetic records. These figures validate the pipeline against known injections; they are not measured plant impact.

  • PostgreSQL
  • FastAPI
  • Next.js
  • SQL
  • Validation
Manufacturing Intelligence Platform executive dashboard with OEE, yield, defects, and downtime metrics
Evaluation appendix containing 21 measured comparisons
Evaluation appendix: 21 measured comparisons

Every figure in these case studies that was measured against a baseline. The units differ and are not comparable to each other; the only thing all 21 share is which way the result came out.

For the build
12
Against it
8
No separation
1
Comparisons measured against a baseline, with the direction of each result
Project Compared Figure Unit Direction
arXiv Recommender Hybrid blend against TF-IDF 0.199 MAP@10 For
arXiv Recommender Hybrid over TF-IDF on MAP@10 1.34× Ratio For
arXiv Recommender MiniLM neural tower against TF-IDF 0.121 vs 0.149 MAP@10 Against
Automotive Analyst Adversarial SQL suite against the guardrail 43 / 43 Attacks blocked For
Automotive Analyst Legitimate analytical queries against the guardrail 16 / 16 Queries accepted For
Automotive Analyst Quoted identifiers against my own allow-list 1 bypass Bypasses found Against
Black Box AI Hybrid retrieval against semantic retrieval 0.902 vs 0.853 MRR No separation
Dispatch Optimizer Greedy plus 2-opt against the naive baseline 12 fewer SLA breaches per day For
Dispatch Optimizer Greedy plus 2-opt against the naive baseline 3.35 hrs Overtime per day For
Dispatch Optimizer Cold CP-SAT at the shipped 8s budget against greedy plus 2-opt 6 behind SLA breaches per day Against
Dispatch Optimizer Warm-started CP-SAT at 8s against greedy plus 2-opt Parity SLA breaches per day Against
Dispatch Optimizer The naive baseline measured against itself 45.6% Deadlines missed Against
Grid Intelligence Model against the published EIA day-ahead forecast 5.49% vs 8.30% MAPE For
Grid Intelligence Model against the published EIA day-ahead forecast +10.0% RMSE skill For
Grid Intelligence Model against EIA, counted per authority 35 / 48 Balancing authorities For
Manufacturing Intelligence Platform Injected process events against what the pipeline recovered Round trip Signal recovery For
Manufacturing Intelligence Platform Propagation join against assuming the origin station 69% Defects downstream For
Manufacturing Intelligence Platform Rediscovered step change against the seeded one 43% Defect reduction For
Water System Risk & Funding Priority Index Weighted index against counting prior violations 0.740 vs 0.734 ROC AUC Against
Water System Risk & Funding Priority Index Compliance component alone against the full index 0.781 ROC AUC Against
Water System Risk & Funding Priority Index Index stripped of compliance against chance 0.447 ROC AUC Against