AI Lab Tech Engineer

Department Icon Data Science Analytics & Machine Learning
149+ Applicants
Posted: 1 week ago
0-15 years
Fremont, CA
hybrid

Posted: 1 week ago
|
Applicants: 149+
Job Description
About Company
Similar Jobs
Please verify your account first! Send OTP

Job Description

Role: AI Lab Tech EngineerLocation-Fremont, CA (Local Candidate Onsite/Hybrid work)
Job Type-Long Term Contract

Were looking for an Infrastructure Engineer to own the execution layer beneath our RL environments: the systems that let an agent operate inside a realistic, multi-tool world coherently for hours or days.

This is a hard systems problem disguised as an AI job.
As the tasks agents can complete keep lengthening, the environments that train them have to stay coherent across far longer horizons than anything that exists today.
That means sandboxing and isolation you can trust, execution thats fast and cheap enough to run at training scale, and the ability to snapshot, restore, inspect, and branch a running environment instead of treating every rollout as one-shot. Youll build the platform that makes all of this possible.

Youll work closely with our research and data teams, and directly with frontier labs and enterprise customers, to turn environment designs into infrastructure that runs reliably in production.

What Youll Do

  1. Environment Execution & Sandboxing:
    • Design and own the sandboxing and execution layer that environments run inside. Build systems to snapshot and restore environment state (disk, process, and where relevant memory and accelerator state) so runs can be paused, resumed, inspected, and branched rather than executed once.
    • Develop the machinery to detect failure modes early in a rollout (reward hacks, infra faults, fairness issues) and to revert to a known-good state, patch, and continue.
    • Extend execution to long-horizon and multi-node environments, where an agent operates across many tools and services over hours or days.
  1. Performance & Scale
    • Own the performance characteristics of the platform: throughput, latency, and cost-per-rollout at scale.
    • Drive utilization and scheduling so we can run far more environment rollouts per dollar without sacrificing reliability.
    • Profile and remove bottlenecks across the stack, from container startup to environment teardown.
    • Build the observability that lets us understand whats happening inside thousands of concurrent, long-running rollouts.
  1. Environment Platform
    • Build and maintain the framework for specifying, packaging, and deploying RL environments which is used by both humans and agents authoring environments internally.
    • Create the tooling that lets researchers and environment authors debug a specific failure across hundreds of long agent traces.
    • Deploy large / small models on on-prem hardware
  1. Collaboration & Production Excellence
    • Scale prototypes into production systems with reproducible workflows and high engineering standards.
    • Write the documentation and tools that let internal teams and external users build on the platform.

What Were Looking For

  1. Systems & Infrastructure

    Looking to get Placed? Try our Placement Guarantee Plan

    • Strong track record building production systems or research infrastructure at scale: distributed systems, execution engines, container/sandboxing infrastructure, or similar.
    • Deep comfort with the systems layer: containers and isolation (e.g. namespaces, cgroups, VMs, gVisor/Firecracker-style sandboxing), filesystems, process and state management.
    • Experience making systems fast and cheap profiling, scheduling, resource utilization, and cost optimization at scale.
    • Proficiency with cloud platforms (Google Cloud Platform, AWS) and distributed computing.
    • Strong engineering fundamentals and a systematic approach to testing, validation, and reliability.
    • Experience of deploying large / small models on on-prem hardware
  1. Execution & Ownership
    • Comfort operating in ambiguity.
    • Strong Python skills; comfort in a systems language (Rust, Go, or C++) is a plus.
    • Ability to use modern tools such as Claude Code effectively.
  1. Collaboration & Communication
    • Excellent communication skills for working with research teams and enterprise customers.
    • Ability to translate between research needs and infrastructure requirements.
    • Comfortable presenting technical work to diverse audiences.

Skills

PythonAiGoogle Cloud

If a job posting appears fraudulent, asks for payment, contains misleading information, or violates our guidelines, please report it immediately. Our team will review it promptly, Jobaaj does not charge any fee from the applicants.

About Company

Akshaya was founded in 1995 under the stewardship of Mr. T. Chitty Babu, a leading light of the real estate industry in India. Since its inception, Akshaya has set the highest standards for itself, inspired by the Sanskrit origin of its name, which means 'endless pursuit'. It has grown into one of the most awarded companies in the country acclaimed for its transparent business practices and innovation. The journey over the last Two Decades has seen the company excel in both the home and commercial domains by building more than 155 magnificent edifices in South India. Powered by a team of more than 200 top-notch professionals and driven by the core values of 'Passion, Principles and Performance', Akshaya has now earned the respect of peers in the construction industry and the admiration of thousands of delighted customers.

Read More

Important dates & deadlines?

Application Deadline

26 Oct 26, 04:15 PM IST

Similar Jobs

View All
Loading...
Bag Logo
Jobaaj
Don't Miss out any Updates

Subscribe now for the latest job alerts
and never miss an update

Job Alert
Google hiring for Specific Roles Apply Now!
1 min ago
New Opportunity
Amazon is hiring freshers Apply Now!
5 min ago
Featured Jobs
Microsoft opening 50+ positions Apply Now!
10 min ago

AI Lab Tech Engineer

Share with