JOB DETAILS
Senior Platform Engineer
CompanyC-Infinity
LocationUnited States
Work ModeOn Site
PostedSeptember 20, 2026

About The Company
No description available for this Company.
About the Role
About C-Infinity
C-Infinity builds foundational AI for mechanical design and manufacturing: models that reason about 3D geometry, motion, spatial constraints, physical feasibility and production logic. AutoAssembler turns that reasoning into production-ready assembly planning inside customer CAD and PLM environments.
Planning runs are the product's unit of work. You own the platform behavior around those runs: APIs, job orchestration, data flow, tenant isolation, queueing, event streams, scaling, observability and the product surfaces that make work visible to customers and to us.
This role is for someone who wants to build the durable middle of a technical product: the systems that let complex planning jobs run predictably as assemblies, tenants and integrations grow.
What you'll do
Job system and event flow
Design and improve the systems that schedule, execute and report on planning runs: queueing, Kafka/Redpanda event streams, per-tenant concurrency and priority, GPU-aware placement, idempotency, retries, dead-letter handling and in-flight work during deploys.
APIs and product platform
Build APIs and internal contracts that make planning jobs understandable and controllable. Expose state, timing, failure reasons, outputs and cost in a way engineers and customers can trust.
Scaling and reliability
Find the parts of the platform that will break as assemblies get larger, runs get longer and customer count grows. Ship the smallest changes that improve throughput, backpressure, latency, durability or operational clarity.
Data stores and integrations
Work across MongoDB/DocumentDB, object storage, CAD and PLM integrations, and customer deployment constraints. Preserve correctness without making every customer environment a custom platform.
Observability and debugging
Own instrumentation for platform behavior: per-job state, queue depth, tenant concurrency, event lag, retries, cost and failure reasons. When a run fails, the system should help explain where, why and what can happen next.
What we're looking for
Seven or more years in backend, platform, infrastructure or SRE roles, with most of the following:
Distributed job and queue systems - you've designed or substantially reworked one (Temporal, Argo, Celery, AWS Batch or similar): scheduling, concurrency limits, retries, idempotency, dead-letter handling and instrumentation.
Kafka, Redpanda or a similar event streaming system used in production.
API design for long-running work, async state, tenant-scoped resources or internal platform contracts.
Production accountability - systems where downtime, data loss or unclear failures had consequences.
Scaling judgment - you can distinguish a future bottleneck from an imaginary one, then make the real one visible.
Working knowledge of cloud infrastructure on AWS or Azure, containers and deployment constraints.
Python or TypeScript for backend/platform work.
Product judgment - you can look at a workflow and say what would make it better for the person using it.
You keep going when the logs don't explain it.
You can explain a design decision to an engineer and to a customer's security reviewer in the same week.
Nice to have
MongoDB/DocumentDB operations and schema evolution.
Temporal, Argo, Celery, AWS Batch or similar workflow/job systems.
GPU workload scheduling or capacity management.
Deploying software into customer-controlled accounts.
CAD or PLM exposure.
Government and regulated deployments
Some of our work runs in US government and other regulated environments. None of this is required; each is a plus.
NIST 800-171 or CMMC Level 2 - or SOC 2, ISO 27001 or FedRAMP taken through to evidence, SSP and POA&M.
AWS GovCloud or Azure Government - including where they diverge from commercial regions.
Export-controlled data - ITAR or EAR-aware handling and access control.
US citizenship, or clearance eligibility - some customer environments are restricted to US persons.
How we work
We remove moving parts before we add them. Where two designs work, we take the one that's simpler at 3am.
Engineering owns product outcomes. If you see a better way for the product to behave, say so - and usually build it.
Portability and simplicity pull against each other. We make that trade deliberately, not by default.
We use AI heavily and expect you to use it well - owning the output as if you'd written every line, because shipping it means you did.
Adopting something proven is a legitimate outcome. Not every problem needs a system built for it.
Flag being out of depth early rather than late.
Details
Hybrid in the Bay Area, plus equity. Ranges for other locations are adjusted to the local market.
C-Infinity is an equal opportunity employer. We consider all applicants without regard to race, color, religion, sex, gender identity, gender expression, sexual orientation, national origin, age, disability, veteran status, or any other protected characteristic under state or federal law. If you need an accommodation at any point in the process, tell us and we will arrange it.
Interested?
Send us a note with your resume or LinkedIn profile and a few lines about the work you want to own.
Apply
C-Infinity builds foundational AI for mechanical design and manufacturing: models that reason about 3D geometry, motion, spatial constraints, physical feasibility and production logic. AutoAssembler turns that reasoning into production-ready assembly planning inside customer CAD and PLM environments.
Planning runs are the product's unit of work. You own the platform behavior around those runs: APIs, job orchestration, data flow, tenant isolation, queueing, event streams, scaling, observability and the product surfaces that make work visible to customers and to us.
This role is for someone who wants to build the durable middle of a technical product: the systems that let complex planning jobs run predictably as assemblies, tenants and integrations grow.
What you'll do
Job system and event flow
Design and improve the systems that schedule, execute and report on planning runs: queueing, Kafka/Redpanda event streams, per-tenant concurrency and priority, GPU-aware placement, idempotency, retries, dead-letter handling and in-flight work during deploys.
APIs and product platform
Build APIs and internal contracts that make planning jobs understandable and controllable. Expose state, timing, failure reasons, outputs and cost in a way engineers and customers can trust.
Scaling and reliability
Find the parts of the platform that will break as assemblies get larger, runs get longer and customer count grows. Ship the smallest changes that improve throughput, backpressure, latency, durability or operational clarity.
Data stores and integrations
Work across MongoDB/DocumentDB, object storage, CAD and PLM integrations, and customer deployment constraints. Preserve correctness without making every customer environment a custom platform.
Observability and debugging
Own instrumentation for platform behavior: per-job state, queue depth, tenant concurrency, event lag, retries, cost and failure reasons. When a run fails, the system should help explain where, why and what can happen next.
What we're looking for
Seven or more years in backend, platform, infrastructure or SRE roles, with most of the following:
Distributed job and queue systems - you've designed or substantially reworked one (Temporal, Argo, Celery, AWS Batch or similar): scheduling, concurrency limits, retries, idempotency, dead-letter handling and instrumentation.
Kafka, Redpanda or a similar event streaming system used in production.
API design for long-running work, async state, tenant-scoped resources or internal platform contracts.
Production accountability - systems where downtime, data loss or unclear failures had consequences.
Scaling judgment - you can distinguish a future bottleneck from an imaginary one, then make the real one visible.
Working knowledge of cloud infrastructure on AWS or Azure, containers and deployment constraints.
Python or TypeScript for backend/platform work.
Product judgment - you can look at a workflow and say what would make it better for the person using it.
You keep going when the logs don't explain it.
You can explain a design decision to an engineer and to a customer's security reviewer in the same week.
Nice to have
MongoDB/DocumentDB operations and schema evolution.
Temporal, Argo, Celery, AWS Batch or similar workflow/job systems.
GPU workload scheduling or capacity management.
Deploying software into customer-controlled accounts.
CAD or PLM exposure.
Government and regulated deployments
Some of our work runs in US government and other regulated environments. None of this is required; each is a plus.
NIST 800-171 or CMMC Level 2 - or SOC 2, ISO 27001 or FedRAMP taken through to evidence, SSP and POA&M.
AWS GovCloud or Azure Government - including where they diverge from commercial regions.
Export-controlled data - ITAR or EAR-aware handling and access control.
US citizenship, or clearance eligibility - some customer environments are restricted to US persons.
How we work
We remove moving parts before we add them. Where two designs work, we take the one that's simpler at 3am.
Engineering owns product outcomes. If you see a better way for the product to behave, say so - and usually build it.
Portability and simplicity pull against each other. We make that trade deliberately, not by default.
We use AI heavily and expect you to use it well - owning the output as if you'd written every line, because shipping it means you did.
Adopting something proven is a legitimate outcome. Not every problem needs a system built for it.
Flag being out of depth early rather than late.
Details
Hybrid in the Bay Area, plus equity. Ranges for other locations are adjusted to the local market.
C-Infinity is an equal opportunity employer. We consider all applicants without regard to race, color, religion, sex, gender identity, gender expression, sexual orientation, national origin, age, disability, veteran status, or any other protected characteristic under state or federal law. If you need an accommodation at any point in the process, tell us and we will arrange it.
Interested?
Send us a note with your resume or LinkedIn profile and a few lines about the work you want to own.
Apply
Key Skills
Distributed SystemsJob Orchestration and QueueingKafka/Redpanda Event StreamingAPI DesignAWS or Azure Cloud InfrastructureContainers and DeploymentPythonTypeScriptMongoDB/DocumentDBObservability and Monitoring
Categories
EngineeringTechnologySoftware Development
Benefits
Equity
Job Information
📋Core Responsibilities
Own the platform behavior around planning runs: job scheduling and orchestration, Kafka/Redpanda event streams, per-tenant concurrency, GPU-aware placement, retries and dead-letter handling, plus the APIs and observability that expose job state, timing, failure reasons and cost. Improve throughput, backpressure, latency and durability as assemblies, tenants and integrations scale, working across MongoDB/DocumentDB, object storage and CAD/PLM integrations.
📋Job Type
full time
📊Experience Level
5-10
💼Company Size
Not specified
📊Visa Sponsorship
No
💼Language
English
🏢Working Hours
Not specified
Apply Now →
You'll be redirected to
the company's application page