Site Reliability Engineer (AI Infrastructure)
Posted bythe hiring team· 17 days ago
- Location
Posted bythe hiring team· 17 days ago
Site Reliability Engineer (AI Infrastructure)
USD 79,622 – USD 95,546
Mid-range for Data
Be among the first applicants
Verified team
HR-vetted before going live.
Transparent pay
Salary stated upfront.
Be among the first applicants
Just opened — your application stands out.
About this role
Key Responsibilities:
Building and maintaining observability for AI workloads, including telemetry, dashboards, alerts, SLO/SLI tracking, and driving improvements when targets are missed.
Writing automation and tooling to reduce operational toil, improve deployment safety, and accelerate incident response.
Integrating AI workloads into existing incident management processes, building runbooks, participating in on-call rotations, and conducting blameless post-mortems.
Building and maintaining CI/CD integrations, deployment safety checks, and rollback automation.
Collaborating with product engineering teams to improve reliability, contribute to architecture decisions, and ensure operational readiness for product releases.
Contributing to capacity planning, autoscaling configuration, and workload scheduling for AI compute infrastructure.
Requirements:
Expertise in SRE, infrastructure, or platform engineering, managing large-scale distributed systems with extensive operational experience.
Expertise in Kubernetes and large-scale containerization systems.
Experience defining SLOs and working with observability tools like Prometheus, Grafana, and distributed tracing to enhance system monitoring.
Proficiency in Python or Go for automation, CI/CD pipelines, deployment safety, and infrastructure-as-code like Terraform.
Interest in or experience with AI/ML infrastructure, model serving, or GPU workloads.
Ability to resolve issues independently while maintaining accountability throughout the process.
Accountability for reliability, developing automation and monitoring, and collaborating effectively with engineering teams unfamiliar with SRE practices.
The Site Reliability Engineer (AI Infrastructure) role with the hiring team offers USD 79,622–95,546 per year. Salary information is published as part of every JobRemotely listing so candidates can self-screen before applying.
Yes — the hiring team has marked this Site Reliability Engineer (AI Infrastructure) role as open to candidates based in Poland. Eligibility requirements are surfaced in the JobPosting structured data on the listing.
The hiring team uses the JobRemotely structured hiring pipeline: candidates apply through the listing, complete a paid test task or screening, and only then proceed to interviews. This skips the resume black hole and respects everyone's time.
Similar roles
Hand-picked from the same category.
the hiring team· Canada, USA·Remote·17 days ago
USD 138,800 – USD 192,800
Viewthe hiring team· US·Remote·17 days ago
USD 230,000 – USD 322,000
Viewthe hiring team· US·Remote·17 days ago
USD 230,000 – USD 322,000
Viewthe hiring team· US·Remote·17 days ago
USD 95,100 – USD 169,800
View