AI Systems Engineer - AI Platforms - Manager
Salary not listed
EY · Charlotte, NC · Hybrid
Build · Found · posted 4 days ago
Apply
Opens EY's own application page in a new tab.
Location: Anywhere in Country
At EY, we’re all in to shape your future with confidence.
We’ll help you succeed in a globally connected powerhouse of diverse teams and take your career wherever you want it to go. Join EY and help to build a better working world.
The opportunity
We are seeking AI Systems Engineers to build and operate the foundational substrate that powers EY’s AI-native platform. This role owns the infrastructure and cloud-native platform layers of the Hybrid AI Multi-Environment Runtime (HAI), from bare-metal and GPU infrastructure through Kubernetes, cluster fabric, and multi-tenant scaling. You will be responsible for building and managing EY Fabric environments across cloud, on-prem, edge, and air-gapped targets.
This is the substrate on which EY Agentic AI capabilities run. This role is ideal for a full-stack infrastructure leader who is equally comfortable with bare-metal and GPU systems and production Kubernetes at scale, who treats reliability and portability as non-negotiable in regulated client contexts, and who understands that the substrate is a product in its own right, measured by the velocity, safety, and portability it unlocks for every team building above it.
Your Key Responsibilities
• Own the cluster & cloud-native platform: compute, Kubernetes and scheduling, cluster fabric/networking, multi-tenancy, and distributed compute, as the substrate for Agentic AI workflows and tooling.
• Own the infrastructure foundation: Ubuntu/OS, BMC/bare-metal, DPU architecture, and NVAIE (GPU/Network/DCGM), ensuring the physical and virtual bedrock is provisioned, patched, and production-ready.
• Stand up and manage EY Agentic AI environments across cloud (EKS/AKS/GKE), on-prem AI Factory (RKE2/NVAIE), edge (K3s), and air-gapped deployment modes, maintaining one consistent stack contract across all targets.
• Deliver foundational platform capabilities such as Infrastructure Management, Kubernetes & Scheduling, and Cluster Fabric Management, so downstream runtime, data, and execution services can run safely and consistently.
• Own cluster lifecycle, autoscaling, GPU pooling/virtualization, and multi-tenancy boundaries (vCluster/Crossplane/Karpenter), providing isolated, elastic capacity per tenant and engagement.
• Own secure execution and inference: Ray Serve, vLLM/NIM/Triton, and NVIDIA Dynamo, with sandboxed execution (gVisor/Firecracker for hosted, NVIDIA OpenShell/vNode for on-prem) for isolated, safe model execution.
• Own cognitive and routing: Envoy AI Gateway, semantic routing (vLLM-SR), model/prompt selection, and streaming response handling — directing each request to the right model under the right constraints.
• Collaborate with DevOps Engineers on deployment and delivery of the platform itself: CI/CD/CV (ArgoCD), infrastructure-as-code / GitOps (Helm/OpenTofu), so environments are reproducible and drift-free.
• Own backup, disaster recovery, and cross-environment replication for high availability (Velero, CloudNativePG, Cilium ClusterMesh), along with patching and platform supply-chain hygiene.
• Ensure the substrate is modular and swappable, so components can be replaced without rewriting consumers, minimizing vendor lock-in while preserving the stack contract.
Skills And Attributes For Success
• Deep expertise in cloud-native platform engineering (Kubernetes, networking, multi-tenancy at scale) and low-level infrastructure and systems engineering (bare-metal, GPU, DPU).
• Advanced understanding of how compute, networking, storage, and scheduling interact to form a reliable, portable substrate.
• Comfortable operating across cloud, on-prem, edge, and air-gapped environments simultaneously, with a portability-first mindset.
• Strong command of AI gateways, semantic routing, and model/prompt selection under latency and cost constraints.
• A passion for ensuring reliability, repeatability, and reduction of operational toil through automation and infrastructure-as-code.
• Ability to…
More Build jobs in the Charlotte region
Senior AI Product Manager, Data Platform & Ontologies | Growth & Transformation
$155K–$220K
Red Ventures · Charlotte, NC · Hybrid
Build · Found · today
AI Solutions Lead
$155K–$175K
Amerit Fleet Solutions · Charlotte, NC · Hybrid
Build · Found · today
Data Scientist
Salary not listed
KellyMitchell Group · Charlotte, NC
Build · Found · yesterday
AI & Cloud FinOps Engineer
Salary not listed
WTW · Charlotte, NC
Build · Found · yesterday
Revenue Technology & AI Automation Manager I
Salary not listed
AvidXchange · Charlotte, NC · On-site
Build · Found · yesterday