Your safety comes first — read before you go. Taf4All only lists job offers published by third parties. We are not the employer, we do not conduct these recruitments, and we cannot guarantee what happens once you make contact — you deal directly with the person or company behind the offer, at your own risk.
Job Title: RLHF Specialist Location: Remote (Worldwide)Job Summary: An RLHF Specialist is responsible for improving and aligning AI models using Reinforcement Learning from Human Feedback (RLHF)methodologies. This role focuses on designing, implementing, and optimizing feedback pipelines that enhance model performance, safety, factual accuracy, and alignment with human values. Responsibilities Generate high-quality preference data by comparing multiple model responses and ranking them based on criteria such as helpfulness, honesty, and harmlessness (HHH). Design complex, multi-turn prompts to stress-test model behavior and expose weaknesses in reasoning or safety. Write detailed “chain-of-thought” explanations and rationales to train reward models on why specific responses are superior. Collaborate with Machine Learning Engineers to analyze model failure modes and identify data gaps that, when filled, will improve reinforcement learning outcomes. Develop and iterate on annotation strategies for preference scoring and reinforcement signals, ensuring consistency across a global team. Proactively probe models to identify vulnerabilities, biases, or hallucination patterns, documenting findings for model optimization. Analyze edge cases where the reward model behaves unexpectedly (e. g., over-indexing on verbosity or style over substance). Provide detailed feedback to ML engineers on reward model failure modes and suggest specific data interventions to correct model behavior. Develop and document templated instruction sets for larger annotation teams. Translate complex reinforcement learning concepts into simple, repeatable tasks for junior reviewers, ensuring high-quality data collection at scale. Monitor model performance over time by maintaining a personal test set of prompts. Regularly re-evaluate new model versions against historical benchmarks to track improvements or regressions in reasoning and alignment. Requirements Minimum of 2 years of experience in Data Annotation, Model Evaluation, Computational Linguistics, or Trust and Safety, specifically working with AI/ML training data. Strong proficiency in Python and deep learning frameworks (Py Torch, JAX, or Tensor Flow). Deep understanding of Reinforcement Learning concepts (PPO, Trust Regions, Reward Hacking) and how they apply to language generation. Hands-on experience fine-tuning open-source models (e. g., Llama 2/3, Mistral, gemma) using techniques like Lo RA/QLo RA. Experience working with annotation tools (Label Box, Scale AI, Snorkel) and managing human-in-the-loop workflows. Ability to diagnose why an RL policy collapsed and adjust hyperparameters or reward structure accordingly. Experience with Constitutional AI or Self-Alignment techniques. Contributions to open-source alignment libraries (TRL, Transformer Reinforcement Learning, Axolotl). Experience with cloud Platforms (AWS Sage Maker, GCP Vertex AI).
---
**
[Click the Apply button below to apply, and Create my CV to build a CV tailored to this offer, professionally]
When tailoring your CV for Odixcity Consulting, directly address how you meet Sage with a specific, verifiable achievement. Use numbers and outcomes—don’t just list duties. This shows you understand their priority and proves you can deliver.
Positioning — Your cover letter must answer one question: why YOU for THIS specific role right NOW? Avoid generic templates — one sentence on what you specifically bring beats three generic paragraphs.
Targeted application — Align your CV to this offer's exact keywords and back every claimed skill with a concrete example. A cover letter showing you understand the specific challenges of this role consistently makes the difference.
🎯 Make your application ATS-ready
ATS (Applicant Tracking Systems) are the software recruiters use to automatically filter CVs before any human reads them. Our CV builder is specifically designed to pass these filters — and it takes under 3 minutes.
Taf4All only lists job offers published by third parties. We are not the employer, we do not conduct these recruitments, and we cannot guarantee what happens once you make contact — you deal directly with the person or company behind the offer, at your own risk.
•Never pay any amount of money — for a file, a training, a uniform, or an interview. A real employer never asks the candidate to pay.
•Never send a photo of your ID card, passport, or banking details before you have physically verified the employer exists.
•Always meet in a public place, during the day — never an isolated address, a private home, or a location you cannot verify in advance.
•Tell a relative or friend exactly where you are going, with whom, and at what time — and share your live location if possible.
•Search the company name online before going: a real business has a trace (website, reviews, other employees, an official address).
•A salary that is far above the market rate for the position and the city is a red flag — be extra cautious.
🧠 Profil recherché Job Title: Excel Specialist Location: Remote Job Type: Contract/ full time Job Summary: We are seeking a highly analytical and detail oriented Excel Specialist who will serve as the backbone of our data operations, translating raw numbers into actionable business intelligence. You will work closely
🧠 Profil recherché Key Responsibilities: Previous warehouse work experience is required. Accurately count and verify quantities of goods. Maintain daily 5 S organization and cleanlinessof the warehouse. Support warehouse management and operations
🧠 Profil recherché Job Applications Specialist The applications specialist will have responsibility for ensuring the effective deployment and support of Temenos T 24 and other banking applications in HOPE’s network of microfinance institutions and banks. This is accomplished through enhancing the parameterization of T
🧠 Profil recherché Job Summary 1. Provide professional, timely, and high quality customer service by responding to customer inquiries and resolving complaints effectively. Communicate fluently with customers, suppliers, and internal teams in both English and French. Maintain accurate customer records and prepare repor