← 返回岗位列表美国IT/互联网fulltime

软件工程师,基础设施与可靠性

雇主

crewAI

地点

远程 · 美国

待遇

$面议

工作模式

远程

截止日期

12月12日

🤖 AI 简历匹配评估

检测你的简历与该岗位的匹配度,免费

免费评估

岗位摘要

About CrewAICrewAI is the leading framework and enterprise platform for building and orchestrating multi-agent AI systems, powering 300M+ agent executions per month across thousands of companies.

岗位职责

About CrewAI
CrewAI is the leading framework and enterprise platform for building and orchestrating multi-agent AI systems, powering 300M+ agent executions per month across thousands of companies. The Agent Management Platform is our control plane for deploying, monitoring, governing, and scaling agents in production. This role owns the infrastructure foundation that keeps it reliable, secure, and fast.
The Role
You'll build and operate the platform infrastructure behind CrewAI's cloud and enterprise deployments. You'll work across multiple hyperscalers - AWS, Azure, and GCP. You’ll work on containers, CI/CD, deployment automation, observability, secrets, networking, and runtime reliability. Your job is to make the product and runtime teams faster while making customer’s production environments safer.
This is not a pure DevOps support role. You'll write code, improve systems, design deployment paths, harden production, and build the internal platform that lets CrewAI scale and scale our customer deployments.
What You'll Do
Own and improve the infrastructure that runs CrewAI's platform: AWS, ECS/ECR, Docker, Kubernetes/Helm, networking, secrets, databases, Redis, and related services.
Build and maintain CI/CD pipelines for build, test, image publishing, migrations, environment promotion, rollbacks, and deploy safety.
Improve reliability across cloud and enterprise deployments: health checks, alerting, incident response, capacity planning, recovery paths, and operational runbooks - and own the front-line on-call rotation and its SLAs.
Partner with runtime engineers on Celery/FastAPI/Redis workloads and with product engineers on Rails/Solid Queue/Postgres production behavior.
Manage production observability and telemetry infrastructure: logs, metrics, traces, dashboards, Sentry/OpenTelemetry plumbing, actionable alerts, and telemetry export to customers' own monitoring systems.
Harden security and compliance posture across IAM, workload identity, secrets management, vulnerability scanning, dependency/image hygiene, and least-privilege access.
Build the tooling and automation that lets field engineers and customers run self-hosted installs themselves - Helm charts, environment config, release artifacts, pre-flight checks, and install runbooks - so engineering does fewer hands-on installs over time.
Reduce operational toil by automating recurring workflows and making deployments boring.
Requirements
What We're Looking For
Strong infrastructure/platform engineering experience in production SaaS environments.
Deep practical experience with AWS, Docker, CI/CD, GitHub Actions, and containerized services.
Experience with ECS and/or Kubernetes; Helm experience is a strong plus.
Comfort operating PostgreSQL, Redis, background job systems, queues, and web services in production.
Strong debugging instincts across app, infra, network, deploy, and dependency layers.
Security-minded approach to IAM, secrets, workload identity, vulnerability management, and production access.
Ability to write reliable automation in Python, Ruby, Go, Bash, or similar.
Calm, rigorous approach to incidents, rollbacks, migrations, and production change management.
Bonus
Experience with AI/agent platforms, workflow runtimes, or high-volume async execution systems.
Experience supporting enterprise/self-hosted deployments.
Terraform or other IaC experience.
SRE background: SLOs, incident review, capacity planning, load testing.
Familiarity with Rails, FastAPI, Celery, OpenTelemetry, or multi-service observability.
Originally posted on Himalayas

申请条件

- 具备云平台基础设施的构建和运维经验,熟悉 AWS、Azure、GCP 等多个超大规模云服务商
- 熟练掌握容器技术,包括 Docker、Kubernetes/Helm、ECS/ECR
- 具备 CI/CD 流水线设计与维护能力,涵盖构建、测试、镜像发布、迁移、环境晋升、回滚和部署安全
- 熟悉可观测性工具和遥测基础设施,如日志、指标、链路追踪、仪表盘、Sentry/OpenTelemetry
- 具备生产环境可靠性工程经验,包括健康检查、告警、事件响应、容量规划、恢复路径和运维手册
- 能够编写代码并改进系统,而非仅提供 DevOps 支持
- 熟悉网络、密钥管理、数据库和 Redis 等相关服务
- 具备安全意识,熟悉 IAM、工作负载身份、密钥管理等安全与合规实践
- 有与后端工程师协作的经验,熟悉 Celery/FastAPI/Redis 或 Rails/Solid Queue/Postgres 等工作负载
- 具备 on-call 轮值经验,能够承担一线值班及 SLA 责任
- 有构建内部平台以支持规模化部署的经验
- 能够设计部署路径并加固生产环境

雇主简介

CrewAI 是领先的多智能体 AI 系统框架和企业平台,提供用于部署、监控、治理和扩展 AI 智能体的控制平面,每月支持超过 3 亿次智能体执行。

对这个岗位感兴趣?

该岗位暂未开放在线申请,顾问可为您推荐同类岗位或申请指导

咨询不收取任何费用,顾问将为您推荐合适的岗位与申请方式

申请海外岗位,英文简历符合当地格式规范吗?

AI 自动评估你与该岗位的匹配度,3 分钟出结果

免费评估简历匹配度

数据来源:Himalayas

岗位信息来源于公开渠道,版权归原作者所有