Eucalyptus Is Hiring
We’re looking for an engineer/researcher to build and optimize systems for serving modern machine learning models in production.
You’ll work across LLM inference, GPU infrastructure, and model serving, with a focus on taking open-source models and making them efficient, reliable, and production-ready across different workloads.
Musts
Solid understanding of deep learning and modern LLM architectures.
Experience with deep learning frameworks such as PyTorch.
Experience loading, running, and serving open-source models.
Basic understanding of inference engines such as vLLM.
Experience working with GPU servers and managing GPU workloads.
Experience building, testing, deploying, and operating applications in production.
Eager to learn, research, experiment, and develop new ideas.
Pluses
Strong experience with inference engine internals such as vLLM, SGLang, or similar.
Experience with inference efficiency techniques such as speculative decoding, quantization, batching, and caching.
GPU/kernel programming experience with CUDA, Triton, or CuTe DSL.
Experience with distributed and multi-GPU inference.
Contributions to ML systems, compilers, kernels, inference runtimes, or open-source model serving.
About Eucalyptus:
Eucalyptus is a technology company active in the field of Artificial Intelligence, which has focused on enhancing the efficiency and performance of advanced AI models by providing innovative solutions in computing infrastructure.