Senior Data Engineer (Hadoop, Spark)
⚲ Kraków
15 800 - 24 300 PLN (PERMANENT)
Wymagania
- PySpark
- Scala
- Airflow
- Hadoop
- Spark
- Hive
- Yarn
- ETL
- SQL
- REST
- Unix
- Linux
- Data pipelines
- Git
- GitHub
- Jenkins
- Ansible
- Jira
- Big data
- Elasticsearch (nice to have)
- Java APIs (nice to have)
- DevOps (nice to have)
- Spark Streaming (nice to have)
- Apache Airflow (nice to have)
- Python (nice to have)
- PL/SQL (nice to have)
Opis stanowiska
O projekcie:
What will you do?
You will work as a key member of a technical team alongside Engineers, Data Analysts and Business Analysts, contributing to a collaborative Agile development process while designing, developing and maintaining scalable data solutions in a dynamic DevOps environment.
We offer
- Hybrid work in Krakow (2 office days per week)- Working in a highly experienced and dedicated team- Benefit package tailored to your needs (medical, sport, lunch subsidy, life insurance, etc.)- Online training and certifications- Access to e‑learning platform- Social events
Wymagania:
- Experience with Pyspark or Scala development and design- Experience using scheduling tools such as Airflow- Knowledge of Hadoop ecosystem including Spark, Hive, YARN and ETL frameworks- Strong SQL and RESTful services knowledge- Experience working on Unix or Linux platforms- Hands-on experience building data pipelines using Hadoop components- Experience with Git, GitHub, Jenkins, Ansible and JIRA- Understanding of big data modelling using relational and non-relational techniques- Experience debugging code and communicating findings to development teams
Nice to have
- Experience with Elasticsearch- Experience developing Java APIs- Experience in data ingestion processes- Understanding of cloud design patterns- Exposure to DevOps and Agile methodologies such as Scrum and Kanban- Experience with Spark streaming- Experience with Apache Airflow in production- Experience with Hadoop ecosystem in enterprise environments- Knowledge of Python backend services- Experience with Scala for high performance systems- Experience in data integration and ETL processes- Knowledge of PL/SQL- Experience with Linux and Unix system operations
Codzienne zadania:
- Define and contribute to software design and development using Pyspark
- Automate testing of new and existing components
- Promote development standards through code reviews and mentoring
- Provide production support and troubleshooting
- Implement tools and processes ensuring performance, scalability and monitoring
- Collaborate with Business Analysts to interpret and implement requirements
- Participate in planning, sprint reviews and retrospectives
- Contribute to system architecture and design
What will you do?
You will work as a key member of a technical team alongside Engineers, Data Analysts and Business Analysts, contributing to a collaborative Agile development process while designing, developing and maintaining scalable data solutions in a dynamic DevOps environment.
We offer
- Hybrid work in Krakow (2 office days per week)- Working in a highly experienced and dedicated team- Benefit package tailored to your needs (medical, sport, lunch subsidy, life insurance, etc.)- Online training and certifications- Access to e‑learning platform- Social events
Wymagania:
- Experience with Pyspark or Scala development and design- Experience using scheduling tools such as Airflow- Knowledge of Hadoop ecosystem including Spark, Hive, YARN and ETL frameworks- Strong SQL and RESTful services knowledge- Experience working on Unix or Linux platforms- Hands-on experience building data pipelines using Hadoop components- Experience with Git, GitHub, Jenkins, Ansible and JIRA- Understanding of big data modelling using relational and non-relational techniques- Experience debugging code and communicating findings to development teams
Nice to have
- Experience with Elasticsearch- Experience developing Java APIs- Experience in data ingestion processes- Understanding of cloud design patterns- Exposure to DevOps and Agile methodologies such as Scrum and Kanban- Experience with Spark streaming- Experience with Apache Airflow in production- Experience with Hadoop ecosystem in enterprise environments- Knowledge of Python backend services- Experience with Scala for high performance systems- Experience in data integration and ETL processes- Knowledge of PL/SQL- Experience with Linux and Unix system operations
Codzienne zadania:
- Define and contribute to software design and development using Pyspark
- Automate testing of new and existing components
- Promote development standards through code reviews and mentoring
- Provide production support and troubleshooting
- Implement tools and processes ensuring performance, scalability and monitoring
- Collaborate with Business Analysts to interpret and implement requirements
- Participate in planning, sprint reviews and retrospectives
- Contribute to system architecture and design
🔍 Dekoder Ogłoszenia
🔴
contributing to a collaborative Agile development process
Może oznaczać pracę w zespole, ale też presję na szybkie dostarczanie zmian i częste spotkania.
🔴
dynamic DevOps environment
Wskazuje na szybkie tempo zmian i potencjalnie częste problemy z infrastrukturą lub procesami.
🔴
Benefit package tailored to your needs
Często oznacza standardowy pakiet z kilkoma opcjami do wyboru, a nie faktycznie dopasowany do indywidualnych potrzeb.
🟡
Experience with Pyspark or Scala development and design
Chociaż wymagane jest doświadczenie, nie określa poziomu zaawansowania ani konkretnych projektów, co może oznaczać różne poziomy oczekiwań.
🟡
Hands-on experience building data pipelines using Hadoop components
Podkreśla praktyczne umiejętności, ale nie precyzuje skali ani złożoności tych potoków danych.