AI Infrastructure
Senior AI Infrastructure Engineer — GPU Clusters & Confidential Computing
- Location
- Menlo Park, CA (San Francisco Bay Area)
- Working arrangement
- Hybrid
- Employment type
- Contractor / Project-based
About Simulation95
Simulation95 is building at the intersection of artificial intelligence, healthcare, and life sciences. We believe meaningful progress starts with people who understand the hardest problems firsthand. We’re bringing together experts, researchers, and builders to turn that understanding into technology that matters.
For curious people who take their craft seriously, this is an opportunity to help shape something early—with challenging problems, meaningful responsibility, and room to make a lasting contribution.
About the role
You will lead the sourcing, design, and deployment of a private AI data center, from hardware procurement and facility planning through secure cluster operation. This is a project-based engagement for an engineer who has built and operated GPU infrastructure.
The planned deployment will grow from 24 to 48 to 104 NVIDIA B300 GPUs, using 3, 6, and ultimately 13 eight-GPU DGX nodes, with high-speed InfiniBand networking and shared high-performance storage. It will support large-model inference, fine-tuning, and training.
Confidential computing and storage are core requirements. You will assess hardware-backed isolation, encryption, remote attestation, and key management across the system, and verify which protections the proposed configuration supports. The architecture must accommodate expansion without replacing its core networking, storage, security, or orchestration systems.
Responsibilities
- Develop the bill of materials, evaluate supplier quotes, and verify compatibility and availability. Manage lead times, warranties, support, and delivery.
- Assess electrical capacity, rack density, power distribution, redundancy, cooling, airflow, floor loading, and connectivity. Coordinate installation with qualified electrical and mechanical specialists.
- Design the initial cluster and expansion plan, accounting for switch ports, bandwidth, storage throughput, rack capacity, and management infrastructure.
- Deploy and tune InfiniBand, RDMA, and shared storage for distributed workloads, model loading, dataset access, and checkpointing.
- Define the threat model and validate CPU/GPU isolation, remote attestation, and protected data paths across compute, interconnects, and storage. Identify unsupported configurations and security gaps before purchase.
- Design encryption and access controls for datasets, model weights, checkpoints, temporary files, and backups. Implement key management, rotation, recovery, and attestation-based key release.
- Deploy Linux, NVIDIA software, containers, scheduling, workload isolation, secure access, and monitoring. Select and configure Kubernetes, Slurm, or an appropriate combination.
- Test performance, security, recovery, and reliability with representative inference, fine-tuning, and training workloads. Document the configuration and operating procedures.
Qualifications
- Direct experience building and operating multi-node NVIDIA GPU clusters using DGX, HGX, or comparable systems.
- Practical experience with InfiniBand, RDMA, NCCL, distributed GPU workloads, and high-performance shared storage.
- Experience implementing hardware-backed confidential computing, CPU/GPU attestation, and secure key management.
- Ability to assess data protection at rest, in transit, and in use, including the boundaries between compute nodes and shared services.
- Strong Linux, virtualization, container, automation, and troubleshooting skills.
- Experience sourcing enterprise hardware and coordinating vendors, integrators, and facility contractors.
- Ability to explain technical tradeoffs, costs, dependencies, and limitations clearly.
Preferred experience
- Direct B200 or B300 deployment experience.
- Intel TDX or AMD SEV-SNP; NVIDIA confidential computing and attestation tooling.
- KMS/HSM integration and confidential virtual machines or containers.
Initial deliverables
The initial work will include a phased architecture and bill of materials, supplier quotes, a facility readiness assessment, an implementation budget and schedule, and a documented security design supported by compatibility evidence. Deployment will include measurable acceptance criteria and a tested expansion plan.