Data Scientist - Product Catalog Data
⚲ Poznań, Wilda
14 600–20 825 zł brutto / mies.
Wymagania
- Spark
- PySpark
- Python
- Apache Airflow
Opis stanowiska
Nasze wymagania:
Hands-on Spark experience - You have experience working with Spark or PySpark. Since we deal with massive datasets, this is a must-have for our daily operations.
An MLOps & deployment mindset - You don't just build models to sit on a laptop - you love seeing them go live. You have practical experience building automated pipelines and successfully deploying models into production environments.
Solid software engineering habits - Python is your go-to language. You write clean, object-oriented code (classes, clean functions, and unit tests) and feel confident navigating and understanding large, complex repositories.
Strong Data Science foundations - You have a deep understanding of statistics and machine learning. You know how to choose the right models, evaluate their performance properly, and calculate sample sizes accurately.
Enthusiasm for GenAI & Agentic AI - You are excited about the future of AI and already use it to supercharge your daily workflow. You have a good theoretical and practical understanding of GenAI principles (prompting, validation) and love using coding assistants (like Copilot) smartly and critically.
Great communication & business sense - You can naturally bridge the gap between tech and business. You are skilled at turning a business challenge into a clear ML problem and explaining complex technical results in a simple, intuitive way to stakeholders.
A relevant background - You hold a degree in a quantitative field (such as Mathematics, Physics, Computer Science, Economics) OR have equivalent, solid hands-on experience in Data Science and Machine Learning roles.
Mile widziane:
Experience with Airflow: If you already know how to use Apache Airflow for pipeline orchestration, that’s a big plus!
O projekcie:
Our team drives the solutions that build a trustworthy marketplace. As a collaborative group of Data Scientists, Data Engineers, and Analysts, we develop the statistical models, products and business rules necessary to analyze and continuously enhance our Product Catalog.
We oversee the quality and correctness of all product components: titles, descriptions, images, and parameters. Our scope also covers ensuring consistency with sellers offers, automated products duplicates detection, building selection models and more. We work in close collaboration with developers and business teams, participating in the entire project lifecycle - from ideation through to implementation. Our ultimate goal is to create a single, reliable source of truth that powers a professional shopping experience for buyers and a fair, frictionless environment for our partners.
Important things for you:
Flexible working hours in the hybrid model (4/1) - working hours start between 7:00 a.m. and 10:00 a.m. We also have 30 days of occasional remote work.
The salary range for this position depending on the skill set is as follows (contract of employment, tax-deductible cost):
- PLN gross 14 600 - 20 825
- Annual bonus based on your annual performance and company results.
Our team is mostly based in Poznan and Warsaw.
Zakres obowiązków:
Co-creating projects on each stage - from concept to productization - to deliver insights and models solving actual business problems
Using a wide range of model modeling techniques, including gradient boosting, Bayesian methods, causal inference, optimization, deep learning, and Generative AI (LLMs, Agentic AI)
Working hand-in-hand with our Data Engineers who support the heavy-lifting of data processing on GCP
Working closely with Analysts, PMs, and Developers to ensure the solutions you build integrate seamlessly into broader initiatives
Leveraging diverse data types, moving beyond tabular data to extract value from spatial data, natural language (NLP), images, and time series
Taking part in implementations of both offline and online models
Engaging in brainstorming sessions, active knowledge sharing, and continuous professional development
Oferujemy:
Flexible working hours in the hybrid model (4/1) - working hours start between 7:00 a.m. and 10:00 a.m. We also have 30 days of occasional remote work.
Long term discretionary incentive plan based on Allegro.eu shares (restricted stock units).
Annual bonus based on your annual performance and company results.
Well-located offices (with e.g. fully equipped kitchens, bicycle parking, terraces full of greenery) and excellent work tools (e.g., raised desks, ergonomic chairs, interactive conference rooms).
A 16" or 14" MacBook Pro or corresponding Dell with Windows (if you don't like Macs) and all the necessary accessories.
A wide selection of fringe benefits in a cafeteria plan - you choose what you like (e.g., medical, sports or lunch packages, insurance, purchase vouchers).
English classes that we pay for related to the specific nature of your job.
A training budget, inter-team tourism (see more here), hackathons, and an internal learning platform where you will find multiple trainings.
An additional day off for volunteering, which you can use alone, with a team, or with a larger group of people connected by a common goal.
Social events for Allegro people - Spin Kilometers, Family Day, Fat Thursday, Advent of Code, and many other occasions we enjoy.
Hands-on Spark experience - You have experience working with Spark or PySpark. Since we deal with massive datasets, this is a must-have for our daily operations.
An MLOps & deployment mindset - You don't just build models to sit on a laptop - you love seeing them go live. You have practical experience building automated pipelines and successfully deploying models into production environments.
Solid software engineering habits - Python is your go-to language. You write clean, object-oriented code (classes, clean functions, and unit tests) and feel confident navigating and understanding large, complex repositories.
Strong Data Science foundations - You have a deep understanding of statistics and machine learning. You know how to choose the right models, evaluate their performance properly, and calculate sample sizes accurately.
Enthusiasm for GenAI & Agentic AI - You are excited about the future of AI and already use it to supercharge your daily workflow. You have a good theoretical and practical understanding of GenAI principles (prompting, validation) and love using coding assistants (like Copilot) smartly and critically.
Great communication & business sense - You can naturally bridge the gap between tech and business. You are skilled at turning a business challenge into a clear ML problem and explaining complex technical results in a simple, intuitive way to stakeholders.
A relevant background - You hold a degree in a quantitative field (such as Mathematics, Physics, Computer Science, Economics) OR have equivalent, solid hands-on experience in Data Science and Machine Learning roles.
Mile widziane:
Experience with Airflow: If you already know how to use Apache Airflow for pipeline orchestration, that’s a big plus!
O projekcie:
Our team drives the solutions that build a trustworthy marketplace. As a collaborative group of Data Scientists, Data Engineers, and Analysts, we develop the statistical models, products and business rules necessary to analyze and continuously enhance our Product Catalog.
We oversee the quality and correctness of all product components: titles, descriptions, images, and parameters. Our scope also covers ensuring consistency with sellers offers, automated products duplicates detection, building selection models and more. We work in close collaboration with developers and business teams, participating in the entire project lifecycle - from ideation through to implementation. Our ultimate goal is to create a single, reliable source of truth that powers a professional shopping experience for buyers and a fair, frictionless environment for our partners.
Important things for you:
Flexible working hours in the hybrid model (4/1) - working hours start between 7:00 a.m. and 10:00 a.m. We also have 30 days of occasional remote work.
The salary range for this position depending on the skill set is as follows (contract of employment, tax-deductible cost):
- PLN gross 14 600 - 20 825
- Annual bonus based on your annual performance and company results.
Our team is mostly based in Poznan and Warsaw.
Zakres obowiązków:
Co-creating projects on each stage - from concept to productization - to deliver insights and models solving actual business problems
Using a wide range of model modeling techniques, including gradient boosting, Bayesian methods, causal inference, optimization, deep learning, and Generative AI (LLMs, Agentic AI)
Working hand-in-hand with our Data Engineers who support the heavy-lifting of data processing on GCP
Working closely with Analysts, PMs, and Developers to ensure the solutions you build integrate seamlessly into broader initiatives
Leveraging diverse data types, moving beyond tabular data to extract value from spatial data, natural language (NLP), images, and time series
Taking part in implementations of both offline and online models
Engaging in brainstorming sessions, active knowledge sharing, and continuous professional development
Oferujemy:
Flexible working hours in the hybrid model (4/1) - working hours start between 7:00 a.m. and 10:00 a.m. We also have 30 days of occasional remote work.
Long term discretionary incentive plan based on Allegro.eu shares (restricted stock units).
Annual bonus based on your annual performance and company results.
Well-located offices (with e.g. fully equipped kitchens, bicycle parking, terraces full of greenery) and excellent work tools (e.g., raised desks, ergonomic chairs, interactive conference rooms).
A 16" or 14" MacBook Pro or corresponding Dell with Windows (if you don't like Macs) and all the necessary accessories.
A wide selection of fringe benefits in a cafeteria plan - you choose what you like (e.g., medical, sports or lunch packages, insurance, purchase vouchers).
English classes that we pay for related to the specific nature of your job.
A training budget, inter-team tourism (see more here), hackathons, and an internal learning platform where you will find multiple trainings.
An additional day off for volunteering, which you can use alone, with a team, or with a larger group of people connected by a common goal.
Social events for Allegro people - Spin Kilometers, Family Day, Fat Thursday, Advent of Code, and many other occasions we enjoy.
🔍 Dekoder Ogłoszenia
🔴
MLOps & deployment mindset - You don't just build models to sit on a laptop - you love seeing them go live. You have practical experience building automated pipelines and successfully deploying models into production environments.
Oczekują, że będziesz nie tylko tworzyć modele, ale także odpowiadać za ich wdrożenie i utrzymanie w środowisku produkcyjnym, co może oznaczać więcej odpowiedzialności niż tylko czysta Data Science.
🔴
Solid software engineering habits - Python is your go-to language. You write clean, object-oriented code (classes, clean functions, and unit tests) and feel confident navigating and understanding large, complex repositories.
Oprócz modelowania, będziesz musiał pisać produkcyjny kod, który jest dobrze przetestowany i łatwy do utrzymania w dużym projekcie.
🔴
Enthusiasm for GenAI & Agentic AI - You are excited about the future of AI and already use it to supercharge your daily workflow. You have a good theoretical and practical understanding of GenAI principles (prompting, validation) and love using coding assistants (like Copilot) smartly and critically.
Oczekują aktywnego wykorzystania i eksperymentowania z najnowszymi technologiami AI, co może wymagać ciągłego uczenia się i adaptacji do szybko zmieniającego się krajobrazu.
🔴
Great communication & business sense - You can naturally bridge the gap between tech and business. You are skilled at turning a business challenge into a clear ML problem and explaining complex technical results in a simple, intuitive way to stakeholders.
Poza aspektami technicznymi, będziesz musiał aktywnie komunikować się z nietechnicznymi interesariuszami i przekładać ich potrzeby na rozwiązania ML.