Pracuj.pl Praca zdalna Mid

Data Engineer (Python & Java)

COGNIZANT

⚲ Kraków, Prądnik Czerwony

18 000 – 26 000 zł / mies.

Wymagania

  • Python
  • Java
  • Kotlin
  • Apache Spark
  • Hadoop
  • Kafka
  • Hive
  • Airflow
  • SQL
  • Docker
  • Kubernetes
  • AWS
  • Azure
  • Google Cloud Platform

Opis stanowiska

Nasze wymagania:
Bachelor's degree in Computer Science, Information Technology, Engineering, or a related field, or equivalent practical experience.
Proven experience as a Data Engineer or in a similar data-focused engineering role.
Strong programming skills in Python, Java, and/or Kotlin.
Hands-on experience with Apache Spark, Hadoop, Kafka, Hive, Airflow, and SQL Server.
Strong understanding of distributed systems, data architecture, and scalable data processing frameworks.
Experience with software testing methodologies and automation frameworks such as PyTest or JUnit.
Knowledge of containerization and orchestration technologies, including Docker and Kubernetes.
Familiarity with modern CI/CD practices and tools.
Strong analytical, problem-solving, and troubleshooting capabilities.
Excellent communication skills and ability to work effectively in a collaborative, cross-functional environment.

Mile widziane:
Experience working with cloud platforms (AWS, Azure, or GCP).
Exposure to data lake, lakehouse, or modern data platform architectures.
Experience with Infrastructure as Code (Terraform, CloudFormation, or similar tools).
Knowledge of data security, governance, and regulatory compliance frameworks.
Experience supporting machine learning or advanced analytics workloads.

O projekcie:
We are looking for a talented and motivated Data Engineer to join our growing data team. In this role, you will design, build, and maintain scalable, high-performance data platforms and pipelines that support critical business operations and analytics initiatives. You will work closely with data scientists, analysts, architects, and cross-functional stakeholders to deliver reliable, secure, and efficient data solutions in a modern, cloud-native environment.

Zakres obowiązków:
Design, develop, and maintain scalable batch and real-time data pipelines using modern data engineering technologies.
Build and optimize data processing solutions leveraging Apache Spark, Kafka, and distributed computing frameworks.
Develop and manage data warehousing solutions using Apache Hive and SQL-based technologies.
Implement and maintain workflow orchestration using Apache Airflow, ensuring reliability and operational excellence.
Establish and enforce data quality, governance, security, and compliance standards across data platforms.
Monitor, troubleshoot, and optimize data infrastructure, pipelines, and processing jobs to improve performance and scalability.
Design and execute automated testing strategies, including unit and integration testing for data workflows.
Develop, deploy, and manage containerized applications and data services using Docker and Kubernetes.
Build and maintain CI/CD pipelines to enable automated testing, deployment, and monitoring of data solutions.
Collaborate with data scientists, analysts, product teams, and business stakeholders to understand requirements and deliver impactful data products.
Document technical solutions, best practices, and operational procedures to support maintainability and knowledge sharing.
Drive continuous improvement initiatives focused on reliability, scalability, and operational efficiency.

Oferujemy:
Work with modern data technologies and large-scale distributed systems.
Collaborate with highly skilled engineers and data professionals.
Contribute to building data products that drive business decisions and innovation.
Opportunities for professional growth, certification, and continuous learning in data engineering and cloud technologies.

🔍 Dekoder Ogłoszenia

🔴
equivalent practical experience
Może oznaczać, że formalne wykształcenie nie jest kluczowe, ale równie dobrze może być próbą obniżenia wymagań dla kandydatów bez dyplomu.
🟡
Strong programming skills in Python, Java, and/or Kotlin.
Choć wymieniono kilka języków, 'i/lub' sugeruje, że biegłość w jednym z nich może być wystarczająca, a niekoniecznie we wszystkich.
🔴
Proven experience as a Data Engineer or in a similar data-focused engineering role.
Określenie 'podobna rola' jest nieprecyzyjne i może obejmować stanowiska, które nie są stricte inżynierią danych.
🟡
Excellent communication skills and ability to work effectively in a collaborative, cross-functional environment.
Często używany zwrot, który może oznaczać potrzebę ciągłej interakcji z różnymi zespołami, co może być czasochłonne.
🔴
Exposure to data lake, lakehouse, or modern data platform architectures.
Słowo 'exposure' sugeruje jedynie podstawową znajomość, a niekoniecznie głębokie doświadczenie w tych obszarach.