← 返回岗位列表新加坡IT/互联网fulltime

技术团队成员(数据智能)

雇主

Reka

地点

远程 · 新加坡

待遇

S$面议

工作模式

远程

🤖 AI 简历匹配评估

检测你的简历与该岗位的匹配度,免费

免费评估

岗位摘要

In this role, you’ll work closely with model researchers, data infrastructure engineers, and cross-functional partners to make sure our data is high quality and can be produced at petabyte scale in a reliable, efficient way. From understanding how data choices show up in model behavior, to building processing pipelines and running the compute behind them,…

岗位职责

In this role, you’ll work closely with model researchers, data infrastructure engineers, and cross-functional partners to make sure our data is high quality and can be produced at petabyte scale in a reliable, efficient way. From understanding how data choices show up in model behavior, to building processing pipelines and running the compute behind them, you’ll help ensure our models are trained on the best data we can get.
What you’ll do
Work with model researchers to define what “good data” means for our models, including quality metrics, validation checks, and acceptance thresholds
Explore open source datasets and create internal ones most suitable to build fundamental World Models
Build algorithms for automated data quality assessment, data domain mixtures, and domain adaptation from synthetic to real data.
Track datasets, metadata, provenance, and versions so experiments are reproducible and it’s clear what data went into which training and evaluation runs
Own CI/CD and development tooling for the data stack (GitHub, Python, PyTorch), and automate repetitive workflows to reduce friction
Track and optimize throughput, storage, and compute utilization across pipelines and related assets
What we’re looking for
Strong ML and deep learning fundamentals with experience building and operating large-scale data and/or compute systems
Comfortable moving between research questions and production engineering: you can dig into data, run analyses, and also ship reliable systems
Demonstrated research experience with data compositions, quality, and dataset releases
Ability to design and execute experiments with convincing unbiased outcomes
Practical experience with distributed processing and orchestration (Spark, Ray, Airflow, or equivalents)
Solid Pythonskills, and familiarity with the tooling around modern model training workflows (datasets, checkpoints, experiment tracking)
Strong instincts around data quality: how to measure it, how to monitor it, and how to prevent regressions as things scale
Able to work in a fast-moving environment, prioritize what matters, and communicate clearly with both researchers and engineers
Bonus: experience with large video datasets, dataset curation for training, or building internal tooling for evaluation/analysis in ML environments
Reka's Mission
Reka's mission is to build useful multimodal artificial intelligence and use it to empower organisations and businesses. We are a globally distributed foundation model startup, headquartered in the San Francisco Bay Area, California. Embracing a remote-first approach, our team brings together top talent from around the world. Our founding team, along with many of our team members, has contributed to many of the breakthroughs in AI over the past decade.
Why Reka?
An Elite Team: Collaborate with top-tier engineers, researchers, operators from renowned organizations like Google DeepMind and Facebook AI Research (FAIR) and successful startups, driving innovation in cutting-edge AI technology.
Massive Market Opportunity: Be part of a rapidly growing industry poised to transform multiple sectors globally, offering the chance to make a significant impact.
Mission-Driven Environment: Work alongside a collaborative, mission-focused team dedicated to advancing AI for meaningful applications.
Inclusive and Open Culture: Thrive in an open and inclusive work environment that values diverse perspectives and fosters creativity.
Generous Benefits: Enjoy 5 weeks of paid leave to recharge, comprehensive healthcare benefits including vision and dental, and additional perks that support your well-being.
Visa Support: We provide visa assistance, including H1B and OPT transfers, for US employees to ensure a smooth transition and support your career with us.

申请条件

- 强大的机器学习和深度学习基础,具有构建和运营大规模数据和/或计算系统的经验
- 能够在研究问题和生产工程之间灵活切换:能够深入分析数据、运行分析,并交付可靠的系统
- 具备数据组成、质量和数据集发布方面的研究经验
- 能够设计和执行实验,并获得令人信服的无偏结果
- 具有分布式处理与编排的实践经验(Spark、Ray、Airflow 或同等工具)
- 扎实的 Python 技能,熟悉现代模型训练工作流相关工具(数据集、检查点、实验跟踪)
- 对数据质量有很强的直觉:如何衡量、监控和确保数据质量
- 能够与模型研究人员、数据基础设施工程师和跨职能合作伙伴紧密协作
- 能够定义“好数据”的标准,包括质量指标、验证检查和验收阈值
- 能够探索开源数据集并创建适合构建基础世界模型的内部数据集
- 能够构建自动化数据质量评估、数据域混合以及从合成数据到真实数据域适应的算法
- 能够跟踪数据集、元数据、来源和版本,以确保实验可复现
- 能够负责数据栈的 CI/CD 和开发工具(GitHub、Python、PyTorch),并自动化重复性工作流
- 能够跟踪并优化跨管道及相关资产的吞吐量、存储和计算利用率

雇主简介

Reka is a globally distributed foundation model startup focused on building multimodal artificial intelligence to empower organizations and businesses.

对这个岗位感兴趣?

该岗位暂未开放在线申请,顾问可为您推荐同类岗位或申请指导

咨询不收取任何费用,顾问将为您推荐合适的岗位与申请方式

申请海外岗位,英文简历符合当地格式规范吗?

AI 自动评估你与该岗位的匹配度,3 分钟出结果

免费评估简历匹配度

数据来源:Jobicy

岗位信息来源于公开渠道,版权归原作者所有