Senior Software Python Engineer
⚲ Warszawa
20 000 – 35 000 zł / mies.
Wymagania
- Python
- Neo4j
- Data Quality
- Data Engineering
- Pipeline Development
Opis stanowiska
•
Work mode - fully remote.
•
Assignment type: B2B
• Start - 01.09.2026.
• Contract length > 10-12 months + extensions.
• Language - English.
• Industry - pharmaceutical.
• Recruitment process - 2 interviews with the client.
• Workload: Full time.
This role is for a Senior Software Engineer – Knowledge Graph, responsible for maintaining data quality, integration, and development of the knowledge graph platform, a vital asset that consolidates diverse drug discovery and gene biology data.
Responsibilities:
• Own data quality: Define and enforce data validation, provenance tracking, and quality metrics across all data sources in the graph.
• Integrate data sources: Collaborate with internal and external data providers to ingest, normalize, and harmonize heterogeneous biomedical datasets.
• Maintain the knowledge graph: Manage the Neo4j schema, data modeling, and pipeline reliability.
• Prepare data for AI/analytics: Ensure graph data supports AI and analytics use cases, including the Blindspot Analysis.
• Collaborate cross-functionally: Work with research scientists, data scientists, and engineers to align scientific needs with reliable data solutions.
Must Haves:
• Strong Python skills for data engineering and pipeline development.
• Hands-on experience with Neo4j (Cypher, schema design, query optimization).
• Proven experience integrating multiple heterogeneous biomedical data sources.
• Deep understanding of data quality practices in production environments.
• Experience resolving schema, identifier, and consistency issues directly with data providers.
• Familiarity with biomedical standards and identifiers (UniProt, Ensembl, ChEMBL, etc.).
• Strong communication skills with cross-functional scientific and technical teams.
Nice to Haves:
• Experience in pharma, biotech, or academic drug discovery.
• Advanced degree (MSc/PhD) in a relevant field.
• Familiarity with graph ML, embeddings, or search/retrieval concepts.
Work mode - fully remote.
•
Assignment type: B2B
• Start - 01.09.2026.
• Contract length > 10-12 months + extensions.
• Language - English.
• Industry - pharmaceutical.
• Recruitment process - 2 interviews with the client.
• Workload: Full time.
This role is for a Senior Software Engineer – Knowledge Graph, responsible for maintaining data quality, integration, and development of the knowledge graph platform, a vital asset that consolidates diverse drug discovery and gene biology data.
Responsibilities:
• Own data quality: Define and enforce data validation, provenance tracking, and quality metrics across all data sources in the graph.
• Integrate data sources: Collaborate with internal and external data providers to ingest, normalize, and harmonize heterogeneous biomedical datasets.
• Maintain the knowledge graph: Manage the Neo4j schema, data modeling, and pipeline reliability.
• Prepare data for AI/analytics: Ensure graph data supports AI and analytics use cases, including the Blindspot Analysis.
• Collaborate cross-functionally: Work with research scientists, data scientists, and engineers to align scientific needs with reliable data solutions.
Must Haves:
• Strong Python skills for data engineering and pipeline development.
• Hands-on experience with Neo4j (Cypher, schema design, query optimization).
• Proven experience integrating multiple heterogeneous biomedical data sources.
• Deep understanding of data quality practices in production environments.
• Experience resolving schema, identifier, and consistency issues directly with data providers.
• Familiarity with biomedical standards and identifiers (UniProt, Ensembl, ChEMBL, etc.).
• Strong communication skills with cross-functional scientific and technical teams.
Nice to Haves:
• Experience in pharma, biotech, or academic drug discovery.
• Advanced degree (MSc/PhD) in a relevant field.
• Familiarity with graph ML, embeddings, or search/retrieval concepts.
🔍 Dekoder Ogłoszenia
🔴
Contract length > 10-12 months + extensions.
Projekt może być krótszy niż deklarowany, a przedłużenia nie są gwarantowane.
🔴
Own data quality
Będziesz odpowiedzialny za problemy z jakością danych, nawet jeśli nie są one bezpośrednio Twoją winą.
🔴
Collaborate with internal and external data providers
Może to oznaczać konieczność radzenia sobie z trudnymi lub niechętnymi dostawcami danych.
🟡
Prepare data for AI/analytics
Może wymagać pracy nad danymi w sposób, który nie jest w pełni zdefiniowany lub jest eksperymentalny.
🟡
align scientific needs with reliable data solutions
Może oznaczać konieczność tłumaczenia bardzo abstrakcyjnych potrzeb naukowców na konkretne rozwiązania techniczne, co bywa frustrujące.