Senior ML Engineer (Python/Big Data)
Posted bythe hiring team· about 4 hours ago
Posted bythe hiring team· about 4 hours ago
Senior ML Engineer (Python/Big Data)
USD 70,832 – USD 93,600
Mid-range for Data
Be among the first applicants
Verified team
HR-vetted before going live.
Transparent pay
Salary stated upfront.
Be among the first applicants
Just opened — your application stands out.
About this role
the company is a leading European software consulting and engineering company. Our mission is to craft clean code and practical solutions with precision and purpose. We foster a dynamic culture rooted in strong engineering, a sense of ownership, and transparency, empowering professionals to make a substantial impact in the software industry.
About the role
Productionizing and scaling an ML-driven data quality system across the organization. The scope of services involves: building and tuning anomaly-detection and clustering pipelines, pairing classic ML with LLM reasoning to flag and explain issues, collaborating with data producers to fix root causes, and creating as well as maintaining validator models that turn detected anomalies into better future data.
Project
Anomalsky
Project Scope
Our client is a NASDAQ-listed B2B data company powering Go-To-Market strategies with a 360-degree view of every customer, a view whose value depends on the quality of billions of person and company records.
Anomalsky is the ML system built to catch what traditional observability misses: row-level semantic anomalies (e.g., a first_name, title, company_name). Three layers, an ML layer (embeddings + unsupervised clustering) flags suspicious records at scale, an LLM layer removes false positives and explains each cluster, and an optional human-in-the-loop lets domain experts resolve whole clusters at once. The MVP already drove ~40k crucial record corrections in production.
What’s next: the MVP is landing on GCP now. Once it’s operational, the mission is to scale Anomalsky across the entire organization, embedding it into Acquisition pipelines and building a real-time variant that scans data before it reaches customers.
The scope of cooperation covers
Productionizing Anomalsky on GCP and scaling it to operational, organization-wide use.
Evolving the ML / LLM / human-in-the-loop design and the feedback loop that turns expert reviews into reusable knowledge.
Prototyping the low-latency real-time variant.
Integrating Anomalsky into existing workflows, starting with Acquisition.
Tech Stack
Python, Airflow, BigQuery, Snowflake, Spark (Dataproc), Databricks, Iceberg, Starburst, Trino, AWS, GCP, Docker, Terraform, Jenkins, GitHub, Scikit Learn, unsupervised anomaly detection (kNN, Isolation Forest, autoencoders), recursive clustering, classifiers on real + synthetic data, MLflow, LLM-based reasoning.
Project environment
ML and data engineers from the company collaborating with customer data engineers and product management.
What we expect in general
Strong Python and production ML skills, with a proven track record of shipping models into real production pipelines.
Hands-on experience using classic ML to surface data quality issues at scale: unsupervised anomaly detection (kNN, Isolation Forest, autoencoders) and clustering on messy real-world tabular data.
Practical experience pairing classic ML with LLMs: using models to flag suspicious records and LLMs for reasoning, false-positive filtering, and the final verification of anomalies.
Solid data engineering background across the modern stack (Airflow, Spark/Dataproc, BigQuery, Snowflake, Iceberg/Trino) and the production toolchain (GCP, Docker, Terraform, CI, MLflow).
Pragmatic, product-oriented approach focused on incremental value delivery and seamless integration into existing workflows.
Professional fluency in English, enabling smooth technical and business discussions in an international environment.
Seems like lots of expectations, huh? Don’t worry! You don’t have to meet all the requirements.
What matters most is your passion and willingness to develop. Apply and find out!
A few perks of being with us
Building tech community
Flexible hybrid work model
Home office reimbursement
Language lessons
MyBenefit points
Private healthcare
Training Package
Virtusity / in-house training
Access to the above perks is optional and completely voluntary for B2B contractors
The Senior ML Engineer (Python/Big Data) role with the hiring team offers USD 70,832–93,600 per year. Salary information is published as part of every JobRemotely listing so candidates can self-screen before applying.
Yes — the hiring team has marked this Senior ML Engineer (Python/Big Data) role as open to candidates based in Poland. Eligibility requirements are surfaced in the JobPosting structured data on the listing.
The hiring team uses the JobRemotely structured hiring pipeline: candidates apply through the listing, complete a paid test task or screening, and only then proceed to interviews. This skips the resume black hole and respects everyone's time.
Similar roles
Hand-picked from the same category.
the hiring team· Poland·Remote·about 2 hours ago
USD 82,216 – USD 132,810
Viewthe hiring team· Poland·Remote·about 2 hours ago
USD 37,946 – USD 60,713
Viewthe hiring team· Poland·Remote·about 2 hours ago
USD 75,892 – USD 86,010
Viewthe hiring team· Poland·Remote·about 2 hours ago
USD 101,189 – USD 121,426
ViewJob ads say a lot about where engineering is going — if you read enough of them. We went through all 335 remote engineering roles on our board, every one with a published salary, and counted what they actually ask for.