← 返回岗位列表印度IT/互联网fulltime

基础设施 Kubernetes 资深软件工程师

雇主

ServiceNow

地点

远程 · 印度

待遇

₹面议

工作模式

远程

截止日期

12月12日

🤖 AI 简历匹配评估

检测你的简历与该岗位的匹配度,免费

免费评估

岗位摘要

About the role We are hiring a staff software engineer (IC4) for the Telemetry Data Platform team, which carries product usage telemetry for the entire ServiceNow platform on a stack built with Kubernetes, Kafka, and ClickHouse.

岗位职责

About the role
We are hiring a staff software engineer (IC4) for the Telemetry Data Platform team, which carries product usage telemetry for the entire ServiceNow platform on a stack built with Kubernetes, Kafka, and ClickHouse. We are distributed across Israel, India, and the Americas.
This is a full stack role with its center of gravity firmly on the infrastructure side — expect roughly two-thirds of your time on telemetry infrastructure, data pipelines, Kubernetes and DevOps work, and the rest on the services and product surfaces on top. You will be equally comfortable designing a distributed system and owning it in production. If you like following a system from the event that fires to the chart that renders — and you want the pager for it — you will feel at home.
You will lead design and delivery of complex, multi-service features, own technical decisions within a domain, and are fully self-directed — escalating only genuine ambiguities and driving decisions to closure. Of the ten critical skills at this level, incident response management is the only one set at Expert; the rest are Experienced.
What you’ll do
Telemetry infrastructure and data pipelines
Design, build, and own backend services for distributed data streaming and processing that handle high-volume, high-cardinality event data with predictable latency and no silent data loss.
Build and maintain Kafka-based streaming pipelines and the pipeline components that feed ClickHouse.
Own data modeling and query performance in ClickHouse — partitioning, sort keys, materialized views, retention, and the cost curve that comes with all of it.
Partner with product and platform teams to shape requirements for telemetry ingestion and processing, then drive the solutions to production.
Kubernetes, DevOps, and production ownership
Deploy, scale, and operate services in production Kubernetes environments, including Helm-based deployments and CI/CD pipelines that make releases repeatable and safe to roll back.
Own observability for what you build — meaningful metrics, useful logs, real tracing, and alerts that fire on customer impact rather than on noise.
Debug and resolve production incidents independently, participate in on-call, run root-cause analysis, and operate against defined service level objectives.
Full stack engineering and technical leadership
Own the full development lifecycle for your work, and build the backend services and APIs that expose telemetry to internal consumers, product surfaces, and AI agents — including API contracts, data models, and schema migrations.
Contribute to the analytics front end — dashboards, funnels, and exploration tools — with an eye on performance against large result sets, and improve the web and mobile capture SDKs.
Lead the design of complex, multi-service features across team boundaries, drive decisions to closure, and write the design docs and postmortems that outlive the conversation.
Raise the bar through code review, test strategy, and automation coverage, and mentor engineers on the team.
Experience and education requirements
The job profile defines IC4 by scope, independence, and impact rather than years served. The minimums below are the screening bar for this requisition.
Requirement
Mandatory minimum
8+ years of Backend software engineering experience
4+ years of Designing and operating distributed systems
3+ years Kubernetes in production, including Helm and CI/CD
3+ years Kafka or equivalent streaming and data pipelines
3+ years On-call and production ownership
Bachelor’s degree in computer science, software engineering, or a closely related technical field
Required
A master’s degree may offset up to one year of the experience minimum. Equivalent practical experience is considered where technical depth is clearly demonstrable.
Required qualifications
Strong backend development experience in Java, Python, Go or equivalent, used in production at scale.
Hands-on experience building distributed systems for data streaming, processing, and storage.
Production experience deploying and managing services on Kubernetes, including Helm and CI/CD pipelines.
Working knowledge of Kafka or an equivalent messaging and streaming system.
Solid SQL skills and hands-on experience with a columnar or analytical data store, including query optimization and physical data modeling.
Demonstrated depth in incident responseand customer escalations
Enough front-end capability to build and debug a data-heavy interface: JavaScript or TypeScript with React or Angular.
Desired qualifications
Production experience with ClickHouse or another OLAP or columnar store — sharding, replication, materialized views, cost tuning.
Exposure to observability tooling: OpenTelemetry, metrics and logging stacks, alerting platforms.
Experience building product analytics or telemetry platforms, or with stream processing frameworks such as Flink or Spark.
Client-side instrumentation — web or mobile SDK development, event capture, session and consent handling.
AI skills
Artificial intelligence strategy is a critical skill at Experienced proficiency for IC4 — it applies to how you build and to what you build.
Confident use of AI coding assistants such as GitHub Copilot or Cursor, paired with the critical eye to review AI-generated code as rigorously as any other — catching logic errors, security anti-patterns, and missed edge cases rather than treating model output as production-ready.
AI-assisted debugging and trace analysis, and a clear understanding of the data privacy rules governing what goes into a prompt.
Integrating LLM APIs into production features — and knowing when a model is the wrong tool for the job. Working knowledge of RAG pipelines, embedding stores, and agent frameworks; familiarity with the Model Context Protocol is a plus.
Exposing telemetry to AI agents through tool interfaces agents can use safely: predictable schemas, bounded results, clear error semantics, sane cost controls.
Designing safeguards for non-deterministic components — retry logic, circuit breakers, human-in-the-loop checkpoints — and hardening against latency spikes and hallucinated output reaching a customer.
Instrumenting AI features in production and judging their quality from telemetry and evaluation rather than impressions.
Work Personas
We approach our distributed world of work with flexibility and trust. Work personas (flexible, remote, or required in office) are categories that are assigned to ServiceNow employees depending on the nature of their work and their assigned work location. Learn more here. To determine eligibility for a work persona, ServiceNow may confirm the distance between your primary residence and the closest ServiceNow office using a third-party service.
Equal OpportunityEmployer
ServiceNow is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race, color, religion, sex, sexual orientation, national origin, age, disability, gender identity, veteran status, or any other category protected by law. In addition, all qualified applicants with arrest or conviction records will be considered for employment in accordance with legal requirements.
Accommodations
We strive to create an accessible and inclusive experience for all candidates. If you require a reasonable accommodation to complete any part of the application process, or are unable to use this online application and need an alternative method to apply, please contact for assistance.
Export ControlRegulations
For positions requiring access to controlled technology subject to export control regulations, including the U.S. Export Administration Regulations (EAR), ServiceNow may be required to obtain export control approval from government authorities for certain individuals. All employment is contingent upon ServiceNow obtaining any export license or other approval that may be required by relevant export control authorities.
From Fortune. ©2026 Fortune Media IP Limited. All rights reserved. Used under license.
It all started when engineer Fred Luddy wrote code that automated a tedious task for his coworker, Phyllis. She cried tears of joy. That moment inspired Fred to build a company that could do that for everyone—freeing people from busywork so they could focus on meaningful work. Today, ServiceNow is the AI control tower for business reinvention. Our ServiceNow AI platform brings together any AI, any data, and any workflow— helping 85% of the Fortune 500® work smarter, faster, and better. We're building an AI-native culture where technology and talent are unstoppable together. And we're just getting started.
Join us to put AI to work for people.
Originally posted on Himalayas

申请条件

- 具备全栈开发能力,重心在基础设施侧
- 能够设计、构建并负责分布式数据流和处理后端服务,处理高容量、高基数事件数据,保证可预测延迟且无静默数据丢失
- 具备构建和维护基于Kafka的流式管道及相关组件(用于向ClickHouse供数)的经验
- 具备ClickHouse数据建模和查询性能优化能力,包括分区、排序键、物化视图、保留策略及成本控制
- 能够与产品和平台团队合作,明确遥测数据摄取和处理需求,并推动方案上线
- 具备在生产Kubernetes环境中部署、扩展和运维服务的能力,包括基于Helm的部署和CI/CD流水线
- 能够领导复杂、多服务功能的设计和交付
- 能够在某一领域内自主做出技术决策,完全自我驱动,仅对真正的模糊点进行上报,并推动决策闭环
- 具备事故响应管理能力(专家级)
- 其余九项关键技能达到经验丰富水平
- 能够设计分布式系统并在生产环境中负责其运维
- 愿意承担所构建系统的on-call职责

雇主简介

ServiceNow是一家全球市场领导者,提供AI增强技术平台,服务超过8100家客户,包括85%的财富500强企业。

对这个岗位感兴趣?

上传简历,AI 免费评估匹配度,一键发送至雇主

无需注册,上传简历即可投递

申请海外岗位,英文简历符合当地格式规范吗?

AI 自动评估你与该岗位的匹配度,3 分钟出结果

免费评估简历匹配度

数据来源:Himalayas

岗位信息来源于公开渠道,版权归原作者所有