Senior Software Engineer, Backend (AI Agent Runtime)
The role involves designing, building, and maintaining low-latency serving stacks for ML models, as well as automating training pipelines and optimizing performance at scale. Additionally, the engineer will create reusable SDKs and mentor others on production ML best practices.
- Remote
- United States
- Posted Aug 3, 2026
- 1 position
Job summary
Cresta unlocks the true potential of the customer experience, turning every conversation into a competitive advantage. Cresta’s unified AI platform combines conversational AI agents, real-time human agent augmentation, and comprehensive conversation intelligence to drive revenue and efficiency gains across every channel. The world’s leading companies, including United Airlines, Cox Communications, and Marriott, use Cresta to power world-class customer experiences every day. Born from the Stanford AI Lab, Cresta has raised more than $270 million from the world’s leading investors, including a16z, Greylock, and Sequoia. Cresta’s leadership includes some of the leading minds in AI today. Our CEO, Ping Wu [https://www.linkedin.com/in/pingwu/], founded and led Google's Contact Center AI and Vertex AI platforms before joining Cresta to build the future of AI-driven customer experiences. Over the next few years, AI is going to redefine how people all over the world interact with businesses every day. Come build that future at Cresta. ABOUT THE ROLE * Build real-time AI agent infrastructure: Design and operate the stateful, low-latency runtime that powers voice and chat AI agents — from LLM streaming and conversation state management to graceful recovery and multi-channel support. * Solve distributed systems problems: Own session management across scaled-out workers — including affinity, checkpointing, crash recovery, and consistency under concurrent access. * Build a function execution platform: Own a serverless-style runtime where customers deploy custom logic — build orchestration, container lifecycle, autoscaling, and versioned rollouts. * Own developer experience and test infrastructure: Build CLI tools, local development environments, and test execution frameworks that let engineers iterate quickly and ship with confidence. * Raise the bar on production quality: Drive observability, incident response, and engineering best practices across the team. What we’re looking for * 5+ years of software engineering experience, with meaningful time spent on infrastructure, platform, or systems work. * Strong Python and Go — both are core to this role, not one primary and one secondary. * Deep understanding of distributed systems: consistency, fault tolerance, state management, concurrency. * Experience with Kubernetes and cloud-native infrastructure. * Experience building developer-facing tooling — CLIs, SDKs, local dev environments, or internal platforms. * Strong communicator who can drive technical decisions, write clear design docs, and mentor others. * High bar for code quality — thorough testing, thoughtful code review, and sustainable engineering practices. * Comfort operating what you build — on-call, incident response, and production ownership. * AI-native workflow — you actively use LLMs and AI-assisted tools in your daily development, and can leverage them to move faster and tackle problems that would otherwise be impractical. Nice-to-haves * Experience with real-time voice or streaming media systems. * Hands-on with LLM integration — streaming inference, prompt orchestration, retrieval-augmented generation. * Experience building serverless or function-as-a-service platforms. * Workflow engines (Temporal, Argo, Airflow) for durable, long-running processes. * Experience in conversational AI or speech domains. * Infrastructure-as-code and GitOps workflows. Perks & Benefits: * We offer Cresta employees a variety of medical, dental, and vision plans, designed to fit you and your family’s needs * Paid parental leave to support you and your family * Monthly Health & Wellness allowance * Work from home office stipend to help you succeed in a remote environment * Lunch reimbursement for in-office employees * PTO: 3 weeks in Canada Compensation for this position includes a base salary, equity, and a variety of benefits. Actual base salaries will be based on candidate-specific factors, including experience, skillset, and location, and local minimum pay requirements as applicable. We are actively hiring for this role in the US and Canada. Your recruiter can provide further details. We have noticed a rise in recruiting impersonations across the industry, where scammers attempt to access candidates' personal and financial information through fake interviews and offers. All Cresta recruiting email communications will always come from the @cresta.ai domain. Any outreach claiming to be from Cresta via other sources should be ignored. If you are uncertain whether you have been contacted by an official Cresta employee, reach out to recruiting@cresta.ai [recruiting@cresra.ai]
What you’ll do
The role involves designing, building, and maintaining low-latency serving stacks for ML models, as well as automating training pipelines and optimizing performance at scale. Additionally, the engineer will create reusable SDKs and mentor others on production ML best practices.
Requirements
Candidates should have over 5 years of experience in production software, with at least 2 years focused on ML platforms or infrastructure. Proficiency in Python and knowledge of Golang, Kubernetes, and serving frameworks are essential.
Benefits
• Comprehensive Medical Coverage • Dental Coverage • Vision Coverage • Flexible PTO • Paid Parental Leave • Retirement Savings Plan • Remote Work Setup Budget • Monthly Wellness Stipend • Monthly Communication Stipend • In-office Meal Program • Commuter Benefits
Listed skills
- KubernetesPreferred
- GoPreferred
- TerraformPreferred
- PythonPreferred
Other relevant skills
Identified from the job description. Confirm important requirements above.
- Python
- Golang
- Kubernetes
- MLOps
- Distributed Systems
- Networking
- Container Security
- Testing
- Code Review
- Continuous Delivery
- Large Language Models
- Real-time Streaming Inference
- Terraform
- Helm
- Speech AI
- Conversational AI
Job areas
- Technology
- Software
- Engineering
- Data & Analytics
Additional details
- Minimum experience
- 5+ years
- Posting language
- English
- Working hours
- 40 hours per week
