uvation
Linux Infrastructure Engineer (Bare Metal, Storage & AI Factory Infrastructure)
Job Overview We are seeking a highly experienced Senior Linux Infrastructure Engineer with deep expertise in Linux administration, bare metal infrastructure, enterprise storage, and next-generation AI Factory / GPU infrastructure platforms . This role is focused on designing, deploying, operating, and troubleshooting large-scale Linux-based infrastructure that powers both traditional enterprise workloads and modern AI/ML environments. This is not a DevOps-focused role . We already have a dedicated DevOps team and are looking for an engineer with extensive hands-on experience in Bare Metal as a Service (BMaaS), GPU infrastructure, high-performance storage, data center operations, and enterprise Linux platforms . The ideal candidate will have experience building and managing infrastructure from the hardware layer up, including servers, networking, storage, GPU clusters, and AI-ready platforms. They should be comfortable working with high-performance computing (HPC), AI Factory environments, and large-scale Linux deployments where performance, reliability, and operational excellence are critical. Key Responsibilities & Required Skills Linux & Bare Metal Infrastructure • Expert-level Linux administration (Ubuntu required; Red Hat and SUSE preferred) • Deep expertise in bare metal server deployment, architecture, provisioning, and lifecycle management • Experience operating Bare Metal as a Service (BMaaS) platforms and large-scale infrastructure environments • Strong understanding of server hardware, including: • BIOS/UEFI • RAID controllers • Firmware management • iLO/iDRAC/IPMI • NICs and SmartNICs • HBA cards • Hardware diagnostics and troubleshooting • Experience designing, implementing, and supporting enterprise Linux infrastructure at scale AI Factory & GPU Infrastructure • Experience deploying and managing GPU-accelerated infrastructure for AI/ML workloads • Understanding of NVIDIA GPU technologies including: • A100, H100, H200, B200, or equivalent GPU platforms • NVIDIA DGX and OEM GPU servers • GPU provisioning and lifecycle management • GPU monitoring and performance optimization • Knowledge of AI Factory architecture and infrastructure requirements • Experience supporting GPU clusters, AI training environments, and high-performance computing (HPC) workloads • Understanding of: • GPU resource allocation and scheduling • Multi-GPU systems • GPU networking requirements • High-bandwidth, low-latency infrastructure design • Familiarity with NVIDIA ecosystem technologies such as: • CUDA • NCCL • GPUDirect Storage • NVIDIA Fabric Manager • NVIDIA Base Command (preferred) Enterprise Storage & Data Platforms • Advanced Linux storage administration: • LVM • XFS, EXT4 • NFS • iSCSI • Fibre Channel SAN • Multipath I/O • Strong hands-on experience with Ceph , including: • Cluster architecture • MON, OSD, MDS • RBD, CephFS, RGW • Capacity planning • Performance tuning • Failure recovery • Experience with high-performance AI storage platforms such as: • WEKA • VAST Data • Dell PowerScale • Pure Storage FlashBlade • NetApp • Understanding of: • NVMe-over-Fabrics (NVMe-oF) • RDMA • GPUDirect Storage • Parallel file systems • AI data pipelines Networking & Infrastructure • Strong networking knowledge: • Bonding • VLANs • Routing • MTU optimization • DNS • DHCP • Experience with high-performance data center networking: • 100G/200G/400G Ethernet • RoCE • RDMA • Spine-Leaf architectures • Familiarity with NVIDIA Spectrum-X, Mellanox/NVIDIA ConnectX adapters, or equivalent technologies • Strong understanding of Layer 2 and Layer 3 infrastructure design and troubleshooting Operations & Reliability • Experience with high availability, clustering, and disaster recovery • Strong troubleshooting skills across: • Linux operating systems • Hardware platforms • GPU infrastructure • Networking • Enterprise storage • Experience supporting mission-critical production environments • Bash and Python scripting for automation and operational efficiency • Experience creating operational documentation, runbooks, and infrastructure standards Nice to Have • Kubernetes infrastructure (especially AI/ML and GPU integration) • KVM, VMware, OpenShift Virtualization, or similar virtualization platforms • Ansible automation • NVIDIA Base Command Manager • Slurm or HPC workload schedulers • Observability and monitoring platforms (Prometheus, Grafana, OpenTelemetry) • Data Center Infrastructure Management (DCIM) tools • IPAM solutions • AWS, Azure, or hybrid cloud exposure We Are Not Looking For • Candidates whose experience is primarily CI/CD pipeline engineering • Engineers focused mainly on Terraform, GitOps, or application delivery pipelines • Cloud-only administrators with limited bare metal, storage, or hardware experience • Professionals whose primary expertise is software development rather than infrastructure engineering Ideal Candidate Someone who has spent years designing, building, and operating enterprise Linux environments, large-scale bare metal infrastructure, storage platforms, and modern AI Factory environments. The ideal candidate understands how to deploy and manage GPU-enabled infrastructure, BMaaS platforms, enterprise storage, and high-performance networking while solving complex operating system, hardware, storage, and AI infrastructure challenges. DevOps experience is a plus, but deep Linux, infrastructure, storage, BMaaS, and AI Factory expertise is the primary requirement. Originally posted on Himalayas
