AI Infrastructure Engineer (LLM Training & Inference, CUDA) — London, UK
The Opportunity
Intercom's Fin is the AI Customer Agent company on a mission to help businesses deliver perfect customer experiences. Fin is the highest-performing AI customer agent on the market — resolving complex issues end-to-end across every channel, powered by Intercom's own AI models. Nearly 30,000 global businesses use the platform. Now Fin is hiring Senior+ AI Infrastructure Engineers to build the systems that train and serve the next generation of its AI products — from GPU all the way up to a user agent that resolves millions of customer queries a month.
What You'll Do
- Implement and scale training pipelines for large transformer and LLM models — from data ingestion and preprocessing through distributed training and evaluation
- Build and optimize inference services that deliver low-latency, high-reliability experiences, including autoscaling, routing and fallbacks
- Work on GPU-level performance: tuning kernels, improving utilization, and identifying bottlenecks across the training and inference stack
- Collaborate closely with ML scientists to bring cutting-edge training and inference methods to production
- Play an active role in hiring, mentoring and developing other engineers
- Raise the bar for technical standards, reliability and operational excellence across Fin's AI platform
Who We're Looking For
- 5+ years in software engineering with a strong record of shipping high-quality products or platforms
- A track record in model training or inference at scale, or low-level GPU coding (CUDA, Triton) — one is great, multiple is even better
- Degree in Computer Science, Computer Engineering or equivalent strong fundamentals
Why Join
You'll join a small, highly technical team at the cutting edge of modern AI infrastructure — building custom models like Fin Apex that outperform frontier models in customer service tasks.
🚀 Ready to Apply?
Similar Roles

