LH
LLMHire
Browse JobsMarket TrendsNewSalariesTrendsCompaniesPricingBlog

Never Miss an AI Job

Get weekly AI job alerts delivered to your inbox.

Join the AI hiring radar. Unsubscribe anytime.

LH
LLMHire

The AI Labor Market Intelligence Platform. Real-time job data, salary benchmarks, and hiring trends from 170+ companies.

Jobs

  • Browse Jobs
  • Companies
  • Job Alerts
  • Post a Job
  • Pricing

Resources

  • Blog
  • CyberOS.devScan code for vulnerabilities
  • EndOfCoding.comStay ahead with AI news
  • Vibe Coding AcademyLearn skills employers want
  • Vibe Coding Ebook22 chapters, 200+ prompts
  • Video Tutorials@endofcoding on YouTube

Company

  • About
  • Contact
  • Privacy
  • Terms

Contact

  • hello@llmhire.com
  • Get in Touch

© 2026 LLMHire. All rights reserved.

VeriduxLabsBuilt by VeriduxLabs
Back to all jobs
S

Principal Cloud Platform Engineer

SambaNova
Austin, Texas, United States; San Jose, California, United StatesOnsite1 months agovia Greenhouse
full-timeprincipal

About the Role

<div class="content-intro"><p style="text-align: left;"><span style="font-size: 12pt;">SambaNova is a leader in next-generation AI infrastructure, delivering a full-stack inference platform for customers worldwide. At the core of SambaNova's technology is the RDU (Reconfigurable Dataflow Unit) — a chip built on a dataflow architecture rather than the traditional GPU model. Its decode performance is especially strong for agentic workloads like multi-turn agents, code generation, and long-running applications. RDUs are packaged into SambaRack, rack-scale hardware that lets customers deploy state-of-the-art models with better performance, greater energy efficiency, and faster time to value.</span></p></div><h3><strong>About the team</strong></h3> <div> <div> <div> <div> <div> <div> <div> <div> <div> <p>The Cloud Platform team owns the production inferencing service that serves SambaNova's models to customers on RDU accelerators, including capacity planning, deployment, monitoring, and incident response across regions in the United States, Asia, Europe, and Latin America.</p> </div> </div> </div> </div> </div> </div> </div> </div> </div> <h3><strong>About the role</strong></h3> <p>As a Principal Cloud Platform Engineer, you will be specializing in our AI Inferencing Service and will be the guardian of its reliability, performance, and scalability. You will bridge the gap between software development and operations, applying an engineering mindset to solve operational challenges. Your primary focus will be ensuring our inference endpoints have exceptional uptime, low-latency response times, and efficient resource utilization, directly impacting the experience of our customers and the success of our AI products. This role includes participating in a shared on-call rotation to maintain 24/7 service reliability.&nbsp;</p> <h3><strong>Responsibilities&nbsp;</strong></h3> <p>Some of your responsibilities will include:</p> <ul> <li>Shared ownership of the production inferencing service across regions, covering availability, latency, performance, change management, and capacity planning</li> <li>Standing-up and automating AI infrastructure in new regions</li> <li>Participating in a shared primary/secondary on-call rotation, and leading incident response</li> <li>Building monitoring, alerting, and dashboards in Prometheus, Grafana, and Datadog for service health, model latency and throughput, and accelerator utilization</li> <li>Finding and eliminating performance bottlenecks</li> <li>Designing auto-scaling policies that handle variable inference loads</li> <li>Managing cloud and on-prem infrastructure as code in Terraform and Ansible</li> <li>Building CI/CD pipelines that safely deploy new model versions and service updates</li> <li>Forecasting infrastructure needs against the product roadmap and usage trends, and working with finance to manage cloud spend</li> <li>Defining and reporting on SLOs and SLIs for the inferencing platform, using that data to prioritize reliability work</li> </ul> <h3><strong>Required Qualifications</strong></h3> <ul> <li>B.S. in Computer Science, Computer Engineering, or related field</li> <li>5+ years of experience in a Site Reliability Engineering, DevOps</li> <li>Experience supporting a large-scale, customer-facing service in a public cloud environment (AWS, GCP, Azure)</li> <li>Strong programming and scripting skills in languages like Python, Go, Rust, or Java</li> <li>Proven experience with containerization and orchestration technologies (Docker and Kubernetes)</li> <li>Deep understanding of monitoring and observability principles and tools (e.g., Prometheus, Grafana, ELK Stack, Datadog)</li> <li>Experience with Infrastructure as Code (e.g., Terraform, CloudFormation)</li> <li>Experience with CI/CD principles and tools (e.g., Jenkins, GitHub Actions, ArgoCD)</li> <li>Strong Linux/Unix system administration fundamentals</li> </ul> <h3>Preferred Qualifications</h3> <ul> <li>Experience in a hybrid environment bridging cloud and on-premise/data center infrastructure.</li> <li>Direct experience supporting ML/AI inferencing services in production.</li> <li>Familiarity with GPU-accelerated computing and optimizing workloads for NVIDIA GPUs for purposes of mapping to RDUs.</li> <li>Knowledge of model serving frameworks like vLLM, SGLang or Ray.</li> <li>Understanding of MLOps principles and practices.</li> <li>Experience with managing and tuning databases (SQL or NoSQL) and caching systems (Redis, Memcached)</li> </ul><div class="content-pay-transparency"><div class="pay-input"><div class="description"><p>Base Salary Range:</p></div><div class="title">Base Pay Range</div><div class="pay-range"><span>$210,000</span><span class="divider">&mdash;</span><span>$280,000 USD</span></div></div></div><div class="content-conclusion"><p><strong><span style="font-family: helvetica, arial, sans-serif; font-size: 12pt;">Submission Guidelines<br></span></strong><span style="font-family: helvetica, arial, sans-serif; font-size: 12pt;">Please note that in order to be considered an applicant for any position at SambaNova Systems, you must submit an application form for each position for which you believe you are qualified.&nbsp;</span></p> <p><strong><span style="font-family: helvetica, arial, sans-serif; font-size: 12pt;">EEO Policy<br></span></strong><span style="font-family: helvetica, arial, sans-serif; font-size: 12pt;">SambaNova Systems is an Equal Opportunity/Affirmative Action Employer. All qualified applicants will receive consideration for employment without regard basis of age (40 and over), color, disability, gender identity, genetic information, marital status, military or veteran status, national origin/ancestry, race, religion, creed, sex (including pregnancy, childbirth, breastfeeding), sexual orientation, and any other applicable status protected by federal, state, or local laws.</span></p> <p style="text-align: left;"><span style="font-family: helvetica, arial, sans-serif; font-size: 12pt;"><strong>Benefits Summary for US-Based, Full-Time Employment Positions</strong><br>SambaNova offers a competitive total rewards package, including the base salary, plus equity and benefits. We cover 95% premium coverage for employee medical insurance, and 77% premium coverage for dependents and offer a Health Savings Account (HSA) with employer contribution. We also offer Dental, Vision, Short/Long term Disability, Basic Life, Voluntary Life, and AD&amp;D insurance plans in addition to Flexible Spending Account (FSA) options like Health Care, Limited Purpose, and Dependent Care. Our library of well-being benefits available to you and your dependents includes a full subscription to Headspace, Gympass+ membership with access to physical gyms, One Medical membership, counseling services with an Employee Assistance Program, and much more.</span></p></div>

Required Skills

PythonKubernetesDockerAWSGCPAzureRustScala

About SambaNova

Enterprise AI platform with custom silicon for generative AI.

Visit Company Website

Ready to Apply?

Join SambaNova and work on cutting-edge AI technology