xAI · Post-Training and RL profile

Measure whether a post-training change is real capability.

Mohamed A M Elansary, PhD — multimodel evaluation under uncertainty, RL-adjacent trajectory measurement, and production agent evaluation sets for post-training and RL loops.

Post-training evaluationRL-adjacent measurementUncertainty quantificationProduction agent evals

Evaluation under uncertainty

  • Six-plus years of multimodel, multi-basin forecast experiments across hydroclimates on Linux/HPC.
  • Compared statistical and physically based stacks, quantified uncertainty, and reported regime-dependent failure modes rather than a single flattering score.
  • That is the measurement analogue of asking whether a post-training or RL update improved reasoning, truthfulness, or real-world capability.

Agents and trajectories

  • Production GPT, Claude, and Gemini agent workflows with retrieval, routing, tenant isolation, provenance, and regression evaluation sets at Vertexium.
  • That maps to inspecting whether a tool-using trajectory reflects intended behavior. It is not reward-model training or RLHF.
  • Daily power user of frontier models in production, matching the posting's power-user qualification without claiming that those models were trained here.

Proposed first contribution

For one post-training or RL loop that already matters for reasoning, truthfulness, or real-world capability — the three outcomes the posting names — define intended behavior and a small failure taxonomy: metric movement without a capability change, slice-specific collapse, disagreement with the intended preference, overconfident tool use. Stand up a small evaluation set with provenance, compare simple baselines, attach uncertainty, and write a clear report before expanding the training loop. This is a proposed measurement approach, not a claim of prior RLHF, DPO, reward-model training, or xAI-internal work.

Honest fit boundary

Direct RL and post-training research depth is a stretch. I have not trained reward models, run RLHF or DPO, or claimed xAI-internal or Grok work. I do not invent metrics or safety research. The posting states that relevant post-training or RLHF experience is not required. The credible contribution is evaluation under uncertainty, RL-adjacent measurement, scientific/HPC rigor, and production agent evaluation harnesses.

Role and location

Member of Technical Staff - Post-Training and RL · Palo Alto, California. The live posting body does not state remote or hybrid eligibility. Willing to relocate to Palo Alto with a relocation package. Remote eligibility is not asserted.

Posting compensation: “$180,000 - $600,000 USD”. · Official role posting