Senior Webscraping & Data Engineer
⚲ Warszawa
38 000 - 48 000 PLN (PERMANENT)
Wymagania
- Data engineering
- Web Scraping
- Python
- pandas
- HTML
- JavaScript
- API
- Data pipelines
- Algorithms
- Data structures
- Data quality
- Troubleshooting
- Problem-Solving
- Communication skills
- Collaboration skills
- Alternative data (nice to have)
- Airflow (nice to have)
- Kafka (nice to have)
- Docker (nice to have)
- Kubernetes (nice to have)
- AWS (nice to have)
- LLM (nice to have)
- MCP (nice to have)
- audio/video data (nice to have)
Opis stanowiska
O projekcie:
Role Overview
We are looking for an experienced Senior Webscraping & Data Engineer to build scalable systems for collecting data from diverse online sources.
The role focuses on reliable, high-performance webscraping and transforming public, unstructured information into datasets for research and investment teams. You will work with scraping, APIs, data processing, automation and AI-driven solutions.
Wymagania:
- 6+ years of experience in software or data engineering.- Strong experience with web scraping, crawling and automated data collection.- Excellent Python skills, including pandas.- Strong SQL and database knowledge.- Good understanding of HTML, JavaScript, APIs and web technologies.- Knowledge of networking concepts related to web data collection.- Experience building data pipelines.- Understanding of algorithms, data structures and data quality.- Experience with large-scale or distributed data collection.- Strong troubleshooting and problem-solving skills.- Interest in AI/LLMs and agentic technologies.- Strong communication and collaboration skills.
Nice to Have- Experience with alternative data in finance or research.- Background in data-intensive organizations.- Experience building tools for researchers, analysts, quants or data scientists.- Knowledge of Airflow, Kafka, Docker or Kubernetes.- Experience with AWS or other cloud platforms.- Familiarity with LLM tools, MCPs, agents or autonomous workflows.- Experience with audio/video data.- Knowledge of statistics and data analysis.
Codzienne zadania:
- Design and develop scalable web scraping and data collection systems.
- Build scrapers and crawlers for websites with different structures and technologies.
- Collect data from websites, APIs and dynamic online sources.
- Clean, transform and standardize raw data.
- Improve accuracy, speed, coverage and reliability of scraping solutions.
- Troubleshoot website changes, data quality and performance issues.
- Build automated pipelines for processing collected data.
- Explore AI/ML and agentic technologies for extraction and validation.
- Support infrastructure for large-scale data collection.
- Implement monitoring and data quality controls.
- Work with researchers and analysts to understand data needs.
- Explore new technologies and hard-to-source data.
Role Overview
We are looking for an experienced Senior Webscraping & Data Engineer to build scalable systems for collecting data from diverse online sources.
The role focuses on reliable, high-performance webscraping and transforming public, unstructured information into datasets for research and investment teams. You will work with scraping, APIs, data processing, automation and AI-driven solutions.
Wymagania:
- 6+ years of experience in software or data engineering.- Strong experience with web scraping, crawling and automated data collection.- Excellent Python skills, including pandas.- Strong SQL and database knowledge.- Good understanding of HTML, JavaScript, APIs and web technologies.- Knowledge of networking concepts related to web data collection.- Experience building data pipelines.- Understanding of algorithms, data structures and data quality.- Experience with large-scale or distributed data collection.- Strong troubleshooting and problem-solving skills.- Interest in AI/LLMs and agentic technologies.- Strong communication and collaboration skills.
Nice to Have- Experience with alternative data in finance or research.- Background in data-intensive organizations.- Experience building tools for researchers, analysts, quants or data scientists.- Knowledge of Airflow, Kafka, Docker or Kubernetes.- Experience with AWS or other cloud platforms.- Familiarity with LLM tools, MCPs, agents or autonomous workflows.- Experience with audio/video data.- Knowledge of statistics and data analysis.
Codzienne zadania:
- Design and develop scalable web scraping and data collection systems.
- Build scrapers and crawlers for websites with different structures and technologies.
- Collect data from websites, APIs and dynamic online sources.
- Clean, transform and standardize raw data.
- Improve accuracy, speed, coverage and reliability of scraping solutions.
- Troubleshoot website changes, data quality and performance issues.
- Build automated pipelines for processing collected data.
- Explore AI/ML and agentic technologies for extraction and validation.
- Support infrastructure for large-scale data collection.
- Implement monitoring and data quality controls.
- Work with researchers and analysts to understand data needs.
- Explore new technologies and hard-to-source data.
🔍 Dekoder Ogłoszenia
🟡
AI-driven solutions
Buzzword — może sugerować dodatkowe obowiązki związane z ML/AI poza samym scrapingiem