Principal Site Reliability Engineer (AI Platform Architecture)
Posted bythe hiring team· 17 days ago
- Location
Posted bythe hiring team· 17 days ago
Principal Site Reliability Engineer (AI Platform Architecture)
USD 92,361 – USD 114,655
Mid-range for Data
Be among the first applicants
Verified team
HR-vetted before going live.
Transparent pay
Salary stated upfront.
Be among the first applicants
Just opened — your application stands out.
About this role
Key Responsibilities:
Defining the reliability architecture for AI compute services, including SLO frameworks, fault tolerance patterns, and advanced capacity planning models.
Driving hands-on development of automation and tooling that scales the SRE team's impact and eliminates operational toil.
Designing a comprehensive observability strategy, leveraging existing platforms to build specialized telemetry and GPU-specific monitoring for AI workloads.
Architecting deployment safety standards, including progressive rollouts, canary analysis, and automated rollback processes.
Embedding reliability into the development lifecycle by influencing product engineering architecture and high-level design decisions.
Mentoring and elevating the SRE team through design reviews, code reviews, and hands-on problem-solving.
Requirements:
Extensive experience in SRE or platform engineering, with a proven track record of impact at a principal or staff level.
Deep expertise in Kubernetes, specifically in managing autoscaling, resource scheduling, and orchestration for compute-intensive workloads.
Advanced programming expertise in Python or Go, with experience building production-grade automation and platform services.
Proven ability to influence cross-team technical decisions and elevate technical standards across engineering departments.
Experience or strong technical interest in AI/ML infrastructure, model deployment, and GPU workload optimization.
A system-level approach to designing reliability into innovative platforms while building strong partnerships with product engineering teams.
The Principal Site Reliability Engineer (AI Platform Architecture) role with the hiring team offers USD 92,361–114,655 per year. Salary information is published as part of every JobRemotely listing so candidates can self-screen before applying.
Yes — the hiring team has marked this Principal Site Reliability Engineer (AI Platform Architecture) role as open to candidates based in Poland. Eligibility requirements are surfaced in the JobPosting structured data on the listing.
The hiring team uses the JobRemotely structured hiring pipeline: candidates apply through the listing, complete a paid test task or screening, and only then proceed to interviews. This skips the resume black hole and respects everyone's time.
Similar roles
Hand-picked from the same category.
the hiring team· Canada, USA·Remote·17 days ago
USD 138,800 – USD 192,800
Viewthe hiring team· US·Remote·17 days ago
USD 230,000 – USD 322,000
Viewthe hiring team· US·Remote·17 days ago
USD 230,000 – USD 322,000
Viewthe hiring team· US·Remote·17 days ago
USD 95,100 – USD 169,800
View