Measure whether a post-training change is real capability.
Mohamed A M Elansary, PhD — multimodel evaluation under uncertainty, RL-adjacent trajectory measurement, and production agent evaluation sets for post-training and RL loops.
Evaluation under uncertainty
- Six-plus years of multimodel, multi-basin forecast experiments across hydroclimates on Linux/HPC.
- Compared statistical and physically based stacks, quantified uncertainty, and reported regime-dependent failure modes rather than a single flattering score.
- That is the measurement analogue of asking whether a post-training or RL update improved reasoning, truthfulness, or real-world capability.
Agents and trajectories
- Production GPT, Claude, and Gemini agent workflows with retrieval, routing, tenant isolation, provenance, and regression evaluation sets at Vertexium.
- That maps to inspecting whether a tool-using trajectory reflects intended behavior. It is not reward-model training or RLHF.
- Daily power user of frontier models in production, matching the posting's power-user qualification without claiming that those models were trained here.
Proposed first contribution
For one post-training or RL loop that already matters for reasoning, truthfulness, or real-world capability — the three outcomes the posting names — define intended behavior and a small failure taxonomy: metric movement without a capability change, slice-specific collapse, disagreement with the intended preference, overconfident tool use. Stand up a small evaluation set with provenance, compare simple baselines, attach uncertainty, and write a clear report before expanding the training loop. This is a proposed measurement approach, not a claim of prior RLHF, DPO, reward-model training, or xAI-internal work.
Honest fit boundary
Direct RL and post-training research depth is a stretch. I have not trained reward models, run RLHF or DPO, or claimed xAI-internal or Grok work. I do not invent metrics or safety research. The posting states that relevant post-training or RLHF experience is not required. The credible contribution is evaluation under uncertainty, RL-adjacent measurement, scientific/HPC rigor, and production agent evaluation harnesses.
Role and location
Member of Technical Staff - Post-Training and RL · Palo Alto, California. The live posting body does not state remote or hybrid eligibility. Willing to relocate to Palo Alto with a relocation package. Remote eligibility is not asserted.
Posting compensation: “$180,000 - $600,000 USD”. · Official role posting