Senior Software Engineer, Inference Platform
About this role
Whether an organization is modernizing internal operations, launching AI-powered products, or transforming customer experiences, Aion takes them from concept to production on a single, unified platform.We work differently than most AI companies: our teams deploy alongside our customers, turning production-ready AI into real business outcomes in weeks, not quarters.We’re a fast-growing, VC-backed startup led by founders with a track record of successful exits. With teams across the US, UK, and India, we’re building the next generation of enterprise AI and we’re looking for exceptional people to help us scale.Who You AreYou're a seasoned engineer who has built and scaled high-performance inference systems for AI/ML workloads. . You understand the complexities of serving models at scale latency optimization, resource orchestration, autoscaling dynamics, and production reliability. You've designed distributed systems that handle thousands of requests per second while maintaining sub-second response times and cost efficiency.Experience with Golang is strongly preferred, and exposure to inference engines (vLLM, TGI, TensorRT), containerization, and distributed systems is an added bonus. You take ownership of platform-level decisions, think strategically about performance vs. cost trade-offs, and want your work to power AI inference for thousands of developers globally.You're product-minded, you understand how your technical decisions impact developers using aion's platform and think about the end-to-end user experience. You're a team player comfortable wearing multiple hats one day you're optimizing inference latency, the next you're joining customer calls to understand their deployment challenges, and the day after you're helping with UI/UX, customer success, documentation and product ops.What You'll DoInference Platform Architecture & Core ServicesDesign and build aion's inference service platform the backbone for serving AI models at scale across diverse workloadsOwn and architect core platform components: AI Gateway, Resource Orchestrator, Runtime Engines, and AutoscalerDesign highly modular, scalable, and extensible low-level designs (LLDs) for inference infrastructure componentsLead high-level design discussions, establish architectural patterns, and drive technical decision-making for the inference stackModel Deployment & Lifecycle ManagementUnderstand and optimize the dynamics of model deployment, version upgrades, and rollback strategiesBuild robust deployment pipelines for seamless model updates with zero-downtime deploymentsDesign intelligent routing systems for multi-model serving, A/B testing, and canary deploymentsImplement strategies for efficient GPU utilization and model cold-start optimizationPerformance & Distributed SystemsImplement highly performant and optimized software for low-latency, high-throughput inference servingBuild and debug production-grade code in distributed systems handling real-time AI workloadsOptimize inference pipelines for latency, throughput, batching efficiency, and resource utilizationDesign fault-tolerant systems with graceful degradation and automatic recovery mechanismsObservability & Engineering ExcellenceBuild high-performance telemetry and observability stack for inference metrics, performance tracking, and debuggingImplement comprehensive monitoring for model latency, throughput, error rates, GPU utilization, and cost per inferenceConduct thorough code reviews to maintain code quality, performance standards, and architectural consistencyEstablish engineering best practices for testing, documentation, and production readiness.
What this role is, and what else it is called
Employers in London advertise this kind of work as Software Engineer, Senior Software Engineer, Web Developer and Full Stack Developer too, so it is worth searching more than one wording. It is a senior Software Engineering role in London, advertised as on-site.
About hiring at Aion123
Aion123 has 2 other live technology roles on its careers page, across Software Engineering and Technology Architecture. Of those, 33% are advertised as fully remote and 0% as hybrid. Elsewhere in its adverts Aion123 asks for AWS, Azure and Google Cloud. See all Aion123 roles, salaries and stack.
How this role compares to the market
1,968 live UK roles list Python right now, from 583 employers. The median advertised salary is £80,000; 30% are advertised as remote. Browse Python roles.
554 live UK roles list C++ right now, from 149 employers. The median advertised salary is £55,000; 48% are advertised as remote. Browse C++ roles.
London has 4,322 live IT vacancies across 1,366 employers, with a median advertised salary of £70,000. Browse London roles.
Senior roles make up 1,633 of the live UK IT market and advertise a median of £80,000.
Market figures are today's snapshot of live UK roles on employer careers pages; see hiring trends.
Technologies mentioned
Detected on Aion123's careers page. ProdReady Recruitment lists this vacancy as an aggregator and is not the employer; applications go to the employer's own site. More IT jobs in London.