LLM Evaluation post training

hr GBP 36

OpenListed onFreelancer.com
Hourly

About the project

Senior ML Engineer / Advisor / Technical co-founder: Model Post-Training & Alignment About Bentham Research Grounding Machine Intelligence in the Humanities Bentham is an applied research lab working at the intersection of AI and the humanities. We build doctorate-authored, peer-reviewed evaluation instruments, training environments, and datasets for frontier AI labs, enterprises, and governments. Our core focus spans ethics, moral reasoning, philosophy, political theory, law, theology, and history. We believe expanding model capabilities requires anchoring machine judgment in human wisdom through bottom-up, scholar-led workflows paired with AI-in-the-loop validation. The Role We are seeking an ML Research Engineer or Technical Advisor with hands-on experience in post-training models at a major frontier AI lab. You will bridge our team of PhD humanities scholars and technical alignment workflows, moving Bentham from core methodology pressure-testing into active project execution. You will help design, build, and validate the pipelines that translate complex humanities rubrics into high-yield evaluation benchmarks and post-training datasets. Key Responsibilities Technical Architecture: Translate scholar-authored, rubric-based datasets into machine-readable formats optimized for SFT, RLHF/RLAIF, DPO, and reward modeling. Methodology Pressure-Testing: Scrutinize and refine our evaluation framework to ensure our scholar-led datasets stand up to the technical standards of frontier lab eval teams. Pipeline & Environment Development: Lead the development of pilot training environments and benchmarking tools that test model capabilities beyond traditional STEM domains. Cross-Domain Collaboration: Interface directly with Bentham CEO Marcus Heal and doctorate domain experts to translate qualitative human reasoning into rigorous, verifiable alignment signals. Qualifications Prior experience in model post-training, preference tuning, or evaluation design at a major AI lab (e.g., OpenAI, Anthropic, Google DeepMind, Meta etc). Deep technical understanding of SFT, RLHF/RLAIF, LLM-as-a-judge evaluation frameworks, and psychometric benchmark design. Ability to translate nuanced qualitative criteria (law, philosophy, history) into precise ML feedback loops. Pragmatic, developer-first mindset with experience taking experimental evaluation hypotheses into production-ready pipelines.

Skills required

This job is listed on Freelancer.com. AiZity aggregates listings for discovery only and is not the employer. To bid or apply, use the button in the sidebar.

Similar jobs

Other open projects with overlapping skills and budget type.

More Machine Learning (ML) jobs →