Measure interaction over time.
Mohamed A M Elansary, PhD — evaluation frameworks, statistical analysis, measurement under uncertainty, and production agent evaluation sets.
Scientific measurement
- Designed multi-model forecast comparisons across basins and hydroclimates.
- Quantified uncertainty and validated imperfect USGS, NOAA, and NASA observations.
- Ran reproducible Python, R, Bash, Linux, and HPC workflows.
Production systems
- Builds GPT, Claude, and Gemini agent workflows at Vertexium.
- Maintains regression evaluation sets for production agent behavior.
- Ships retrieval, routing, tenant isolation, provenance, and validation systems.
Proposed evaluation approach
Define intended interactive behavior and a failure taxonomy; build a small evaluation set with provenance and ambiguity labels; implement statistical analysis, stratification, and uncertainty; compare simple baselines; and report what the signal does and does not support before embedding it in a model-development loop.
Honest fit boundary
I have not authored an LLM-as-judge system or a Cartesia-domain audio evaluation stack. I do not claim model-alignment research, RLHF, or generative-audio metric invention. My contribution is evaluation frameworks, statistical analysis under uncertainty, honest failure-mode reporting, and production agent regression evaluation.
Role and location
Researcher, Evals · “*HQ - San Francisco, CA” · OnSite. Relocation with a support package is an honest discussion point; remote or hybrid eligibility is not asserted.
“$220K – $350K • Offers Equity” · Official role posting