The Work
General Machines builds the data infrastructure that makes frontier AI possible. That means pipelines that collect behavioral data at scale, annotation tooling that extracts signal from noise, evaluation systems that measure agent behavior precisely, and the infrastructure that keeps all of it running reliably in production environments.
This role is for someone who wants to own the whole stack. You will build systems that run in live environments, process behavioral traces from AI agents doing real work, and create the infrastructure that turns raw session data into training-ready datasets and evaluation benchmarks. You will not be handed a narrow ticket queue. You will own problems.
What You Will Build
Data collection pipelines. We run AI agents across live e-commerce environments and physical settings, generating behavioral traces at continuous scale. You will build and maintain the systems that capture, structure, and store those traces with high reliability.
Annotation and quality control tooling. Raw traces need to become labeled, training-ready data. You will build the tooling that makes annotation fast, consistent, and auditable, and the quality control systems that catch labeling drift before it reaches the dataset.
Evaluation infrastructure. CartBench and our physical AI benchmarks need to run reliably against new models and agent versions. You will build the evaluation harness, scoring systems, and comparison tooling that makes benchmark results meaningful.
SLM deployment and inference infrastructure. Shelf-1 and Commerce-1 are deployed on client hardware. You will work on the inference stack, packaging, and on-premises deployment systems that make on-site model deployment something a non-technical team can actually operate.
What We Are Looking For
Strong Python. You have built production data pipelines and you understand what it takes to make them reliable, not just functional.
Experience with distributed systems and data engineering at a scale where things break in interesting ways.
Some familiarity with machine learning infrastructure: model packaging, inference serving, or training pipelines. You do not need to be a researcher but you need to understand what researchers need from their infrastructure.
Experience with computer vision, LLMs, or AI agents is genuinely useful here. If you have built something in this space, tell us what it actually did and what broke.
Ability to work with ambiguity. This is an early role at a small company. The scope will be wide and the spec will often be thin. If that sounds frustrating rather than interesting, this is probably not the right fit.
No specific degree requirement. We care about what you have built.
The Role
This is an early engineering hire. You will help define how we build, not just execute against an existing architecture. The decisions you make in the first year will shape the system for a long time. We are a small team by design and intend to stay that way.
Remote. Full-time. We work across time zones and communicate mostly in writing.
Apply
Send a note to founders@generalmachines.ai with a description of something you built that you are proud of and what specifically was hard about it. A resume is welcome but not required as the first thing we read.