ByteDance
ByteDanceVerified Source

Large Language Model Inference Engineer ByteDance

128K–256K
Onsite · San Jose, California
Posted July 27, 2026
payroll

Overview

ByteDance seeks a graduate to engineer a large model inference system at Volcano Ark. You will build and scale high‑performance inference clusters while reducing costs through advanced techniques. This role drives innovation in AI infrastructure and impacts enterprise solutions across multiple ByteDance products.

What You'll Do11

  • 1Design and implement the MaaS inference system for Volcano Ark
  • 2Optimize large model inference performance cost and stability at scale
  • 3Reduce inference costs via disaggregated multi‑role inference and heterogeneous inference
  • 4Apply elastic computing and multi‑tenant co‑located inference strategies
  • 5Collaborate with cross‑functional teams to ship features within 90 days
  • 6Drive research on GPU performance analysis and low‑level system bottlenecks
  • 7Maintain and evolve Kubernetes and Ray based resource orchestration frameworks
  • 8Explore distributed network communication optimization using RDMA principles
  • 9Contribute to Kubernetes and Ray frameworks for scalable workload management
  • 10Support internal teams with custom ML compute and algorithm development
  • 11Promote performance improvements through system‑level engineering practices

Requirements10

  • 1Currently completing or recently graduated with a BS or MS in Computer Science or related field
  • 2Strong grasp of algorithms design patterns data structures operating systems and computer architecture
  • 3Proficiency in C++ or Python with solid coding standards
  • 4Deep understanding of GPU hardware CUDA and performance analysis techniques
  • 5Demonstrated interest in distributed systems large‑scale heterogeneous inference and low‑level performance study
  • 6Experience with PD disaggregation and KV cache systems for multi‑node inference
  • 7Expertise in distributed communication optimization RDMA operator implementation
  • 8Familiarity with Kubernetes and Ray for resource orchestration and scheduling
  • 9Graduate status or recent completion within the past year
  • 10Ability to start onboarding by year‑end

Salary Insight

$128 - $256k per year

Location

Typeonsite
LocationSan Jose, California

Required Skills

C++PythonCUDAGPU Performance AnalysisDistributed SystemsKubernetesRayRDMAKV Cache SystemsLarge Model Inference OptimizationPD Disaggregation
Share:

Similar open positions

Explore active roles that match your skills and interests.

ByteDance

ByteDance

18h agoSan Jose, Californiapayroll

Large Language Model Training System Engineer Graduate Applied Machine Learning

ByteDance seeks a graduate to develop Volcano Ark training systems for large model post-training and reinforcement learning. This role involves designing elastic training solutions and optimizing distributed systems. The ideal candidate thrives in a fast-paced environment and contributes to cutting-edge AI infrastructure.

128K–256K
PythonRustC+++7 more
zoox

zoox

18d agoSan Francisco, Californiapayroll

Machine Learning Engineer Multi-Modality Foundation Model

Design and manage system capacity for perception models on autonomous vehicles. Lead initiatives to share perception stack components for efficient inference. Focus on delivering production-ready large-scale models for on-vehicle stacks. Seek experts experienced in compressing accelerating deploying complex models for power-constrained SoCs.

Competitive salary
C++PythonCUDA+6 more

NUBYT, Inc.

7d agoSan Jose, Californiapayroll

LLM Research Engineer IV at NUBYT Inc Mountain View

We seek a skilled LLM Research Engineer IV to design and fine-tune state-of-the-art Large Language Models. This role drives next-generation generative AI bridging cutting-edge NLP research and scalable production systems. The ideal candidate thrives in a collaborative environment and contributes directly to impactful outcomes.

250K–260K
PyTorchTensorFlowJAX+5 more
Technogen, Inc.

Technogen, Inc.

10d agoCharlotte, North Carolinapayroll

Senior LLM Inference & GPU Systems Engineer

Technogen, Inc. seeks a Senior On-Premise LLM Inference & GPU Systems Engineer to architect and optimize high-performance inference platforms in Charlotte, NC. You will own the deployment and tuning of large language models on on-premise GPU clusters, ensuring low-latency, high-throughput serving for enterprise workloads. Collaborate with data scientists and infrastructure teams to build scalable ML pipelines. This role offers direct impact on production AI systems within a Woman-Owned Small Business with 15+ years of IT services.

687K
nvidia tritonkubernetespython+2 more
Intel

Intel

17h agoSan Jose, Californiapayroll

AI Infrastructure Engineer Intel

Performance‑obsessed AI Infrastructure Engineer at Intel in San Jose, California. You will drive inference performance and redefine peak performance on Intel’s next‑generation GPU architectures.

170K–315K
C++PythonGPU Computing+8 more
Bain & Co.

Bain & Co.

12h agoHouston, Texaspayroll

Senior AI/ML Engineer, LLMOps & RAG Systems

You build production inference, serving, and LLMOps infrastructure for Python-based ML systems, taking models from prototype to governed deployment. Bain's Private Equity Group Innovation team creates proprietary data and software products, and your work supports over 1,000 professionals across the investment lifecycle. You collaborate with Data Scientists, Data Engineers, and the Agent/AI squad to deliver reliable, observable ML systems. This hands-on role sets engineering standards and mentors mid-level engineers, with daily use of MLflow, Kubernetes, and Databricks.

141K–169K
PythonMLflowLLMOps+12 more