NoFluffJobs Stacjonarnie Senior

Senior AI Engineer (Evaluations - Canvas Agent)

Instructure Hungary Ltd

⚲ Budapest

18 055 - 26 481 PLN (PERMANENT)

Wymagania

  • Python
  • LLM
  • Evaluation
  • Automated test
  • Prompt Engineering
  • Agentic Architectures
  • TypeScript (nice to have)
  • Node.js (nice to have)

Opis stanowiska

O projekcie:
At Instructure, we believe in the power of people to grow and succeed throughout their lives. Our goal is to amplify that power by creating intuitive products that simplify learning and personal development, facilitate meaningful relationships, and inspire people to go further in their education and careers. We do this by giving smart, creative, passionate people opportunities to create awesome things.

Our Advanced Development team builds AI-native capabilities, reusable AI systems, and shared infrastructure that power multiple products and workflows across the platform.

We are looking for a Senior Data Scientist to help build, ship, and scale production AI systems from the ground up. This is an engineering-forward ML/AI systems role, intended for candidates who are comfortable moving from prototype to production and can own critical parts of the ML/AI systems lifecycle, including pipelines, model integration, inference services, deployment, monitoring, and operational reliability.

You will work closely with product, engineering, and research partners to turn advanced AI ideas into reliable product capabilities used at scale.

Important note on scope: This role is not primarily focused on BI/reporting or experimentation analytics. We are looking for someone with strong experience building and operating production ML/AI systems.

Growth & Impact - In This Role, You’ll Be Expected To

- Influence architecture and engineering standards for AI systems.
- Shape reusable infrastructure and platform patterns.
- Mentor a growing team of AI/ML engineers.
- Turn advanced AI research into production value for educators and learners at scale.

Wymagania:
- LLM Evaluation Experience: A background that clearly demonstrates hands-on experience designing and writing automated tests for LLMs. You must have a proven track record of writing and implementing LLM-as-a-judge evaluations in real-world scenarios.- Technical Stack: A strong background in Python is highly preferred, or a demonstrated willingness and ability to learn it quickly. You should also have professional experience in full-stack or backend engineering to support both the evaluation infrastructure and general product development needs (experience with TypeScript/Node.js is a plus).- Prompt Engineering Expertise: Deep understanding of how to reliably prompt models for classification, extraction, and grading tasks without falling prey to common biases (e.g., position bias, verbosity bias).- Evaluation Tooling: Familiarity with modern LLM observability and evaluation frameworks (e.g., LangSmith, Braintrust, Ragas, promptfoo, or similar).- Agentic Architectures: Strong conceptual understanding of how AI agents plan and execute tool calls (e.g., ReAct, tool-use APIs) so you can effectively evaluate multi-step workflows.- Analytical Mindset: An ability to translate highly subjective concepts (like "helpfulness" or "tone") into rigorous, trackable metrics.- Collaboration Skills: Ability to work across product and engineering teams to understand the Canvas Agent's use cases and align evaluation criteria with user needs.

Onsite Collaboration Requirement: This role requires working onsite on Tuesday and Wednesday, with Thursday strongly encouraged as part of our company’s in-person collaboration model.

Codzienne zadania:
- Design the Evaluation Framework: Build and maintain scalable LLM-as-a-judge pipelines to automatically score the Canvas Agent’s actions, responses, and tool usage across a variety of complex educational workflows.
- Develop Rubrics & Datasets: Create comprehensive grading rubrics and curate high-quality "golden" datasets (both real and synthetically generated) to baseline and test the agent's performance.
- Optimize Judge Prompts: Engineer and iterate on prompts for the judges, ensuring automated scoring aligns with high quality evaluations.
- Full-Stack Contribution: Step beyond evaluation pipelines to participate in full-stack product development as needed, collaborating with the team to build and refine the core AI features, UI components, and application architecture.
- Accelerate Iteration: Integrate your automated evaluations directly into our CI/CD pipelines, creating a "paved path" that allows our AI product teams to ship updates with high velocity and total confidence.
- Analyze & Report: Monitor evaluation metrics to identify failure modes, hallucination rates, and regressions. Translate these subjective quality signals into objective, actionable engineering tasks.

🔍 Dekoder Ogłoszenia

🔴
This is an engineering-forward ML/AI systems role, intended for candidates who are comfortable moving from prototype to production and can own critical parts of the ML/AI systems lifecycle, including pipelines, model integration, inference services, deployment, monitoring, and operational reliability.
Oczekuje się, że będziesz nie tylko tworzyć modele, ale także zajmować się całym procesem wdrażania i utrzymania ich w środowisku produkcyjnym, co może oznaczać dużą odpowiedzialność i szeroki zakres obowiązków technicznych.
🟡
You will work closely with product, engineering, and research partners to turn advanced AI ideas into reliable product capabilities used at scale.
Może to oznaczać, że będziesz musiał tłumaczyć skomplikowane koncepcje techniczne na język biznesowy i współpracować z różnymi zespołami, co wymaga dobrych umiejętności komunikacyjnych i kompromisów.
🟢
Influence architecture and engineering standards for AI systems.
Masz potencjał do kształtowania przyszłości technologii AI w firmie, ale może to również oznaczać, że będziesz musiał przekonywać innych do swoich pomysłów i radzić sobie z oporem.
🔴
Shape reu
Ten fragment jest niekompletny, co sugeruje, że ogłoszenie może być niedopracowane lub wymagać doprecyzowania.