← 返回岗位列表中国专业服务fulltime

高级站点可靠性工程师

雇主

Playson

地点

远程 · 中国

待遇

面议

工作模式

远程

截止日期

12月13日

🤖 AI 简历匹配评估

检测你的简历与该岗位的匹配度,免费

免费评估

岗位摘要

About the Role We’re looking for a Senior Site Reliability Engineer to join our Infrastructure Squad - a lean & senior team where ownership is high and expectations are even higher. This is a deeply hands-on role at the core of a high-traffic system, where you’ll be directly responsible for maintaining reliability, performance, and stability…

岗位职责

About the Role
We’re looking for a Senior Site Reliability Engineer to join our Infrastructure Squad - a lean & senior team where ownership is high and expectations are even higher. This is a deeply hands-on role at the core of a high-traffic system, where you’ll be directly responsible for maintaining reliability, performance, and stability in a fast-paced environment.
You’ll be working on real-time production challenges, handling incidents, managing alerts, and being part of a critical on-call rotation. This role requires resilience, strong decision-making under pressure, and a proactive mindset to continuously improve systems operating at scale.
If you thrive in high-load environments, enjoy solving complex production issues, and want to have a direct impact on systems used by millions - this is the place for you.
Key Responsibilities
Own system reliability by actively monitoring platform health, managing alerts, and responding to incidents in real time
Participate in 24/7 on-call rotations, taking full ownership of production stability in a high-traffic (5–7k RPS) environment
Investigate incidents, perform root cause analysis, and implement long-term fixes to prevent recurrence
Build and continuously improve monitoring, alerting, and observability across the Kubernetes (EKS) ecosystem
Deploy, manage, and optimise infrastructure using Terraform, Helm, and GitOps tools (Flux/ArgoCD)
Drive automation and proactively improve system resilience, reducing manual intervention and recurring issues
Maintain and evolve CI/CD pipelines and infrastructure-as-code practices
Collaborate closely with engineering teams to support deployments and minimise user impact in a live environment
Introduce and integrate new tools and technologies to enhance scalability, reliability, and performance
Handle environment-specific requests and ensure smooth day-to-day platform operations under constant load
Requirements
Strong hands-on experience with Kubernetes (deployment, scaling, troubleshooting) in high-load environments
Experience with GitOps tools such as FluxCD or ArgoCD
Proven experience in incident response, root cause analysis, and postmortems in production systems
Solid experience with AWS, Terraform, Docker, and CI/CD pipelines
Experience with monitoring and observability tools such as Datadog, Prometheus, Grafana, and logging stacks like ELK or CloudWatch
Strong understanding of networking concepts and protocols
Proficiency in at least one scripting language (e.g. Python, Go, Node.js)
Experience working with version control systems (Git)
Familiarity with incident management tools like PagerDuty, Opsgenie, or similar
Ability to operate effectively in a fast-paced, high-pressure environment with strong ownership and accountability
Proactive, resilient mindset with a focus on continuous improvement and system stability
What We Offer
Competitive Salary
Quarterly Bonuses
Unlimited Paid TimeOff
Unlimited Paid SickLeave
Remote & Flexible Working
Private MedicalInsurance
Financial Supportfor Life Events
Professional DevelopmentBudget
International Exposure
Regular CompanyEvents
*Benefits may vary depending on location and contractual agreement
Recruitment Process
1. HR Interview (30-45 min)
2. Technical interview (90 min)
4. Final Interview with C-level (60 min)
By submitting your application, you acknowledge that your personal data will be processed in accordance with our Privacy Policy.

申请条件

- 具备 Kubernetes 实操经验(部署、扩展、故障排查)
- 熟悉 EKS 生态系统
- 熟练使用 Terraform、Helm 和 GitOps 工具(Flux/ArgoCD)
- 有 CI/CD 流水线和基础设施即代码实践经验
- 具备监控、告警和可观测性工具经验
- 有高流量生产环境工作经验(5–7k RPS)
- 能参与 24/7 待命轮值
- 具备事件响应、根因分析和长期修复能力
- 有自动化经验,能减少人工干预和重复问题
- 能在压力下做出决策,具备韧性
- 有主动改进系统可靠性和性能的心态
- 具备与工程团队紧密协作的能力
- 有引入和集成新工具及技术的经验
- 熟悉实时生产环境中的部署支持和用户影响最小化

雇主简介

Playson is a company that provides IT support and operations, managing employee IT requests, access administration, and endpoint management. The role focuses on internal IT services, suggesting the company operates in the technology sector, possibly in software or gaming.

对这个岗位感兴趣?

该岗位暂未开放在线申请,顾问可为您推荐同类岗位或申请指导

咨询不收取任何费用,顾问将为您推荐合适的岗位与申请方式

投递国内企业,你的简历符合 HR 筛选标准吗?

AI 自动评估你与该岗位的匹配度,3 分钟出结果

免费评估简历匹配度

数据来源:Jobicy

岗位信息来源于公开渠道,版权归原作者所有