Senior Data Scientist – AI/LLM Evaluation
⚲ Remote
23 520 - 28 560 PLN (B2B)
Wymagania
- Data science
- Statistical analysis
- Experimantation skills
- AI
- LLM
- Agentic workflow
- 3D (nice to have)
- Geometry (nice to have)
- Rendering (nice to have)
- LLM-as-a-Judge (nice to have)
- LLM-based evaluation (nice to have)
Opis stanowiska
O projekcie:
We are looking for a Senior Data Scientist to join an international project focused on Artificial Intelligence, Large Language Models (LLMs), and AI evaluation.The project focuses on developing and improving methodologies for evaluating the quality and performance of AI models and agentic workflows. The role involves designing evaluation metrics, building ground-truth datasets, analyzing model performance, and developing advanced approaches to automated and human-feedback-based evaluation.
Wymagania:
- Proven experience as a Senior Data Scientist or in a similar role.- Strong data science, statistical analysis, and experimentation skills.- Hands-on experience with AI/LLM evaluation.- Strong understanding of LLMs and agentic workflows.- Experience designing, validating, and optimizing evaluation metrics for AI/ML models.- Strong analytical skills and experience investigating model errors, including false positives and false negatives.- Experience working with ground-truth datasets and defining quality criteria.- Ability to independently analyze experimental results and translate findings into actionable recommendations.- Strong problem-solving and analytical mindset.Nice to Have- Basic knowledge of 3D, geometry, or rendering.- Experience with LLM-as-a-Judge / LLM-based evaluation systems.- Experience with human-in-the-loop evaluation methodologies.- Previous experience evaluating AI agents or agentic systems.- Familiarity with advanced approaches to evaluating LLM-generated content and outputs.
Codzienne zadania:
- Design and validate evaluation metrics and ground-truth datasets.
- Analyze evaluation quality, including accuracy, false positives, and false negatives.
- Optimize metric normalization and scoring methodologies.
- Develop and improve LLM-based judges for automated evaluation.
- Design and implement human-feedback-based evaluation approaches.
- Explore and develop new evaluation metrics, including:
- complexity,
- prompt adherence,
- output quality and correctness.
- Analyze experiment results and identify opportunities to improve model performance and evaluation methodologies.
- Collaborate with Data Science, AI/ML, and Engineering teams to develop evaluation solutions for AI models and agentic workflows.
We are looking for a Senior Data Scientist to join an international project focused on Artificial Intelligence, Large Language Models (LLMs), and AI evaluation.The project focuses on developing and improving methodologies for evaluating the quality and performance of AI models and agentic workflows. The role involves designing evaluation metrics, building ground-truth datasets, analyzing model performance, and developing advanced approaches to automated and human-feedback-based evaluation.
Wymagania:
- Proven experience as a Senior Data Scientist or in a similar role.- Strong data science, statistical analysis, and experimentation skills.- Hands-on experience with AI/LLM evaluation.- Strong understanding of LLMs and agentic workflows.- Experience designing, validating, and optimizing evaluation metrics for AI/ML models.- Strong analytical skills and experience investigating model errors, including false positives and false negatives.- Experience working with ground-truth datasets and defining quality criteria.- Ability to independently analyze experimental results and translate findings into actionable recommendations.- Strong problem-solving and analytical mindset.Nice to Have- Basic knowledge of 3D, geometry, or rendering.- Experience with LLM-as-a-Judge / LLM-based evaluation systems.- Experience with human-in-the-loop evaluation methodologies.- Previous experience evaluating AI agents or agentic systems.- Familiarity with advanced approaches to evaluating LLM-generated content and outputs.
Codzienne zadania:
- Design and validate evaluation metrics and ground-truth datasets.
- Analyze evaluation quality, including accuracy, false positives, and false negatives.
- Optimize metric normalization and scoring methodologies.
- Develop and improve LLM-based judges for automated evaluation.
- Design and implement human-feedback-based evaluation approaches.
- Explore and develop new evaluation metrics, including:
- complexity,
- prompt adherence,
- output quality and correctness.
- Analyze experiment results and identify opportunities to improve model performance and evaluation methodologies.
- Collaborate with Data Science, AI/ML, and Engineering teams to develop evaluation solutions for AI models and agentic workflows.
🔍 Dekoder Ogłoszenia
🔴
developing and improving methodologies for evaluating the quality and performance of AI models and agentic workflows
Może oznaczać tworzenie od podstaw nowych, skomplikowanych metodologii, lub po prostu dostosowywanie istniejących narzędzi i procesów.
🔴
building ground-truth datasets
Może oznaczać tworzenie od podstaw dużych, wysokiej jakości zbiorów danych, lub po prostu wykorzystywanie i etykietowanie istniejących danych.
🔴
analyzing model performance
Może oznaczać głęboką analizę przyczynową błędów modeli, lub po prostu generowanie standardowych raportów z metryk.
🔴
developing advanced approaches to automated and human-feedback-based evaluation
Może oznaczać tworzenie innowacyjnych, przełomowych metod ewaluacji, lub po prostu implementację znanych technik.
🔴
Ability to independently analyze experimental results and translate findings into actionable recommendations
Oczekuje się samodzielności i proaktywności w wyciąganiu wniosków i proponowaniu rozwiązań, co może oznaczać brak ścisłego wsparcia ze strony przełożonych.