HPC Kubernetes Solutions Architect (GPU Platforms)
Location: Dallas, TX (Hybrid)
Type: Direct Hire
⢠Competitive base salary + performance bonus
⢠100% company-paid benefits
We are seeking an HPC Kubernetes Solutions Architect to lead the design, integration, and adoption of GPU-accelerated Kubernetes platforms supporting HPC, AI/ML, simulation, and scientific workloads.
This is a highly technical, customer-facing architecture role with ownership across the full solution lifecycleâfrom discovery and requirements gathering through architecture design, proof-of-concept delivery, deployment, and long-term optimization. The role serves as a trusted advisor to customers while also influencing internal product and engineering direction through real-world feedback.
The ideal candidate brings deep expertise across Kubernetes, GPU orchestration, and HPC environments, along with the ability to design scalable, high-performance platforms and guide customers through complex infrastructure transformations.
⢠Serve as the primary architectural point of contact for customers adopting GPU-accelerated Kubernetes platforms
⢠Capture workload requirements, performance objectives, and scaling needs, translating them into reference architectures and solution designs
⢠Lead customer workshops, technical design sessions, and architecture reviews
⢠Architect and operate Kubernetes clusters optimized for GPU workloads using NVIDIA GPU Operator, Network Operator, DCGM, and device plugins
⢠Integrate Multi-Instance GPU (MIG), GPU sharing, and advanced scheduling (Volcano, Slurm integration, kube-scheduler plugins)
⢠Design and implement multi-tenant Kubernetes environments with strong isolation and performance guarantees
⢠Develop or extend custom Kubernetes operators and controllers using Go or Python
⢠Automate HPC infrastructure services and platform operations
⢠Support Infrastructure-as-Code and GitOps practices using Terraform, Helm, Kustomize, ArgoCD, and FluxCD
⢠Lead proof-of-concept and benchmarking initiatives to validate performance and scalability
⢠Utilize profiling tools and workload characterization methodologies to optimize GPU utilization and cluster performance
⢠Conduct performance tuning across compute, storage, and networking layers
⢠Define integration strategies across compute, storage, networking, and orchestration layers
⢠Support CNI integrations (NVIDIA CNI, Multus, Cilium), distributed storage (Lustre, GPFS, Ceph, VAST), and container runtimes
⢠Ensure seamless integration with HPC schedulers and enterprise systems
⢠Implement monitoring and telemetry solutions using Prometheus, Grafana, DCGM Exporter, and OpenTelemetry
⢠Provide visibility into GPU health, cluster utilization, and workload performance
⢠Partner with HPC, ML, DevOps, and platform teams to ensure scalability and performance in hybrid and on-prem environments
⢠Collaborate with product and engineering teams to influence roadmap and platform improvements
⢠Build relationships with ecosystem vendors including NVIDIA, networking providers, and storage partners
⢠Stay current on GPU roadmaps, interconnect technologies (InfiniBand, RoCE, NVLink), and Kubernetes advancements
⢠Provide forward-looking guidance to customers on scaling and future architecture evolution
⢠Represent the organization in technical workshops, design sessions, and industry events
⢠Extensive experience designing and operating Kubernetes platforms in HPC or GPU-intensive environments
⢠Deep expertise across:
⢠Proven ability to design scalable, secure, and resilient Kubernetes-based architectures
⢠Proficiency in Go or Python for operator development and automation
⢠Experience with workload profiling, benchmarking, and performance tuning
⢠Strong customer-facing skills with the ability to translate requirements into actionable architectures
⢠Experience collaborating across engineering, product, and operations teams
⢠Experience delivering end-to-end HPC or AI/ML solutions from design through deployment and optimization
⢠Familiarity with containerized HPC environments (e.g., Singularity/Apptainer)
⢠Experience with GitOps practices and CI/CD pipelines for Kubernetes platforms
⢠Contributions to open-source projects in Kubernetes or NVIDIA ecosystems
⢠Experience advising customers on future-state architectures and emerging technologies
⢠Bachelorâs or Masterâs degree in Computer Science, Engineering, Physics, or related field
⢠Relevant certifications such as CKA, CKAD, CKS, AWS Solutions Architect, or Azure Solutions Architect Expert
Job Title: Road Service Tech/Grade A- Mechanic Job Description: Technician is responsible for repair and maintenance of a variety of agricultural, lawn maintenance and construction equipment. This technician may be required to go out into the field and service equipment...
...Connect Staffing is hiring a Batch Maker / Kitchen Operator for a contract food manufacturer in Pittsburgh, PA. This is a temp-to-hire evening position paying $20.60/hr. In this role, you will measure and combine bulk ingredients per production recipes and operate industrial...
...to be part of an inclusive, adaptable, and forward-thinking organization, apply now. We are currently seeking a Utilization Review Nurse - USRN (Remote work!) to join our team in Manila, National Capital Region (PH-00), Philippines (PH). Performs clinical reviews...
...The Landscape Construction and Irrigation Sales Manager oversees the design, estimating, and sales process for commercial and high-end residential landscape and irrigation projects. This role leads a team of designers and estimators to deliver accurate, creative, and...
...Events & Marketing Manager Location: Traders Village Houston Schedule: Full-time, FridayTuesday (weekends mandatory) Are you a creative, organized, and results-driven marketing professional who loves bringing people together through unforgettable events? Traders...