Senior. AI Engineer – LLM Research Engineer
About the role
Senior LLM Research Engineer
About Chubb
Chubb is the world's largest publicly traded property and casualty insurer, with operations in 54 countries providing commercial and personal insurance, reinsurance, and life insurance to clients worldwide.
We're at the forefront of AI-driven transformation in insurance. Chubb is making strategic investments in artificial intelligence to fundamentally change how we assess risk, serve customers, and operate our business. Combining cutting-edge AI capabilities with our exceptional financial strength, comprehensive product portfolio, and global reach, we're building the future of intelligent insurance solutions.
The Role
As a Senior LLM Research Engineer you'll own the loop from dataset to trained checkpoint to measured result. Post-training and inference performance sit in one role because in practice they are one loop: train, quantize, serve, measure, then feed the result into the next run.
What makes it interesting is what you train against. Underwriting, claims, and the other core areas each bring their own vocabulary and their own definition of a correct answer, across diverse language tasks with an emphasis on long-form reasoning and complex instruction following. Ground truth largely does not exist yet: the corpora are large and were never assembled for training, so building training and evaluation sets from them is a novel problem in its own right. We are data-centric by conviction, and post-training data quality is where most of our gains come from.
This is applied science aimed at direct implementation. People on this team move between training, systems engineering, and agentic work as priorities shift. We run closer to a dense model than a mixture of experts.
Major Responsibilities
Run post-training end-to-end: SFT, GRPO and other RL methods, and distillation Design and build evaluation: judge rubrics, eval harnesses, regression suites, and the analysis that turns a training run into a decision Build ground truth where none exists: construct labelled training and evaluation sets from existing corpora, including data synthesis at scale, define what a correct answer looks like per task, and validate that the sets measure what they claim to. Data quality decides whether a run is worth doing Run distributed training on multi-node GPU clusters: parallelism strategy, sharding and offload, memory and throughput tuning, and debugging runs that fail at hour nine Inference performance engineering: tensor and pipeline parallel configuration on vLLM, quantization, speculative decoding, and throughput tuning for both evaluation and serving
What You'll Bring
Deep expertise in production Python, with strong working knowledge of the training and inference stack (PyTorch, Transformers, TRL, DeepSpeed, vLLM, or equivalents) Hands-on post-training experience. You've run SFT and at least one RL method on a real model, and can explain what went wrong the first time Experience with multi-node distributed training: parallelism strategy, sharding and offload, and the failure modes that only appear at scale Evaluation rigour, and the judgement to design measurement that survives contact with a business environment Fluency with agentic coding tools in your own workflow, such as Claude Code and Codex. The team is fully immersed in this way of engineering, and we expect it to be part of how you build rather than something reached for occasionally The profile can come from either direction: a machine learning engineer who has gone deep on training, or an ML researcher whose work has made it to impact stage. Either way we expect engineering discipline, meaning version control, tests, reproducibility, and code the next person can pick up
Strong Preference Given To
Building labelled training and evaluation sets from unlabelled source data, including synthetic data generation LLM-as-judge evaluation design, including inter-rater agreement and rubric validation Inference optimization: quantization, speculative decoding, or kernel-level performance work Domain-heavy language tasks where the correct answer is ambiguous and nuanced
At Chubb we are committed to providing equal employment opportunities to all employees and applicants. It is our policy to provide equal employment opportunities to employees and applicants based on job-related qualifications and ability to perform a job. If you require an accommodation during the hiring process or upon hire, please inform Human Resources. If a selected applicant requests accommodation during the recruitment process, Chubb will consult with the applicant in order to provide suitable accommodation that takes into account the applicant’s accessibility needs.
Not the right fit? Search for Senior. AI Engineer jobs in Toronto, Ontario, Canada
About Chubb
Chubb is a world leader in insurance. With operations in 54 countries and territories, Chubb provides commercial and personal property and casualty insurance, personal accident and supplemental health insurance, reinsurance and life insurance to a diverse group of clients. As an underwriting company, we assess, assume and manage risk with insight and discipline. We service and pay our claims fairly and promptly. The company is also defined by its extensive product and service offerings, broad distribution capabilities, exceptional financial strength and local operations globally. Parent company Chubb Limited is listed on the New York Stock Exchange (NYSE: CB) and is a component of the S&P 500 index. Chubb maintains executive offices in Zurich, New York, London, Paris and other locations, and employs approximately 40,000 people worldwide.
Read our Social Media Guidelines here: https://www.chubb.com/us-en/about-chubb/chubbs-social-media-guidelines.aspx
Notre section « À propos » est également disponible en français, ici: https://www.chubb.com/ca-fr/about-chubb-in-canada/a-propos-de-chubb-au-canada.aspx
Similar Jobs
Senior. AI Engineer – LLM Research Engineer
About the role
Senior LLM Research Engineer
About Chubb
Chubb is the world's largest publicly traded property and casualty insurer, with operations in 54 countries providing commercial and personal insurance, reinsurance, and life insurance to clients worldwide.
We're at the forefront of AI-driven transformation in insurance. Chubb is making strategic investments in artificial intelligence to fundamentally change how we assess risk, serve customers, and operate our business. Combining cutting-edge AI capabilities with our exceptional financial strength, comprehensive product portfolio, and global reach, we're building the future of intelligent insurance solutions.
The Role
As a Senior LLM Research Engineer you'll own the loop from dataset to trained checkpoint to measured result. Post-training and inference performance sit in one role because in practice they are one loop: train, quantize, serve, measure, then feed the result into the next run.
What makes it interesting is what you train against. Underwriting, claims, and the other core areas each bring their own vocabulary and their own definition of a correct answer, across diverse language tasks with an emphasis on long-form reasoning and complex instruction following. Ground truth largely does not exist yet: the corpora are large and were never assembled for training, so building training and evaluation sets from them is a novel problem in its own right. We are data-centric by conviction, and post-training data quality is where most of our gains come from.
This is applied science aimed at direct implementation. People on this team move between training, systems engineering, and agentic work as priorities shift. We run closer to a dense model than a mixture of experts.
Major Responsibilities
Run post-training end-to-end: SFT, GRPO and other RL methods, and distillation Design and build evaluation: judge rubrics, eval harnesses, regression suites, and the analysis that turns a training run into a decision Build ground truth where none exists: construct labelled training and evaluation sets from existing corpora, including data synthesis at scale, define what a correct answer looks like per task, and validate that the sets measure what they claim to. Data quality decides whether a run is worth doing Run distributed training on multi-node GPU clusters: parallelism strategy, sharding and offload, memory and throughput tuning, and debugging runs that fail at hour nine Inference performance engineering: tensor and pipeline parallel configuration on vLLM, quantization, speculative decoding, and throughput tuning for both evaluation and serving
What You'll Bring
Deep expertise in production Python, with strong working knowledge of the training and inference stack (PyTorch, Transformers, TRL, DeepSpeed, vLLM, or equivalents) Hands-on post-training experience. You've run SFT and at least one RL method on a real model, and can explain what went wrong the first time Experience with multi-node distributed training: parallelism strategy, sharding and offload, and the failure modes that only appear at scale Evaluation rigour, and the judgement to design measurement that survives contact with a business environment Fluency with agentic coding tools in your own workflow, such as Claude Code and Codex. The team is fully immersed in this way of engineering, and we expect it to be part of how you build rather than something reached for occasionally The profile can come from either direction: a machine learning engineer who has gone deep on training, or an ML researcher whose work has made it to impact stage. Either way we expect engineering discipline, meaning version control, tests, reproducibility, and code the next person can pick up
Strong Preference Given To
Building labelled training and evaluation sets from unlabelled source data, including synthetic data generation LLM-as-judge evaluation design, including inter-rater agreement and rubric validation Inference optimization: quantization, speculative decoding, or kernel-level performance work Domain-heavy language tasks where the correct answer is ambiguous and nuanced
At Chubb we are committed to providing equal employment opportunities to all employees and applicants. It is our policy to provide equal employment opportunities to employees and applicants based on job-related qualifications and ability to perform a job. If you require an accommodation during the hiring process or upon hire, please inform Human Resources. If a selected applicant requests accommodation during the recruitment process, Chubb will consult with the applicant in order to provide suitable accommodation that takes into account the applicant’s accessibility needs.
Not the right fit? Search for Senior. AI Engineer jobs in Toronto, Ontario, Canada
About Chubb
Chubb is a world leader in insurance. With operations in 54 countries and territories, Chubb provides commercial and personal property and casualty insurance, personal accident and supplemental health insurance, reinsurance and life insurance to a diverse group of clients. As an underwriting company, we assess, assume and manage risk with insight and discipline. We service and pay our claims fairly and promptly. The company is also defined by its extensive product and service offerings, broad distribution capabilities, exceptional financial strength and local operations globally. Parent company Chubb Limited is listed on the New York Stock Exchange (NYSE: CB) and is a component of the S&P 500 index. Chubb maintains executive offices in Zurich, New York, London, Paris and other locations, and employs approximately 40,000 people worldwide.
Read our Social Media Guidelines here: https://www.chubb.com/us-en/about-chubb/chubbs-social-media-guidelines.aspx
Notre section « À propos » est également disponible en français, ici: https://www.chubb.com/ca-fr/about-chubb-in-canada/a-propos-de-chubb-au-canada.aspx