This job is no longer available
This job expired on 23/09/2026. It no longer accepts applications.
LLM Evaluator – Model Response Analyst
Odixcity Consulting
Job description
About the role
We are looking for a detail‑oriented LLM Evaluator to assess and improve the performance of large language models. You will analyze AI‑generated content for accuracy, coherence, factual reliability, bias, safety, and alignment with our guidelines.
Key responsibilities
- Evaluate and rank model‑generated text using complex rubrics covering factuality, coherence, safety, instruction‑following, and creativity.
- Compare multiple model responses to the same prompt and justify the preferred output.
- Provide concise feedback to modeling and training teams about recurring failure patterns.
- Craft adversarial prompts to expose biased, harmful, or insecure outputs and help patch safety gaps.
- Collaborate with QA to refine evaluation guidelines for ambiguous or edge‑case scenarios.
- Participate in cross‑checking sessions to calibrate scoring standards and ensure inter‑rater reliability.
- Investigate underperforming model outputs, hypothesise root causes, and flag novel behaviors for research.
Required profile
- Minimum 2 years of professional experience in computational linguistics, data analysis, technical writing, NLP quality assurance, or cognitive science.
- Bachelor’s degree in computer science or a related field.
- Strong ability to explain why a model output is good or bad based on logic, tone, factuality, and instruction adherence.
- Experience with Reinforcement Learning from Human Feedback (RLHF) data collection and evaluation team coordination.
Required skills
- Prompt engineering
- RLHF data collection
- Dataset sourcing, cleaning and annotation
- Data analysis
- A/B testing concepts
- Experiment design for model comparison
Questions fréquentes
Why are you reporting this job?
Boost your chances
Upload your CV — we will match you with relevant openings.
Analyzing your CV...
Odixcity Consulting