Senior Software Engineer (Agentic Coding RL Environments)
Build coding RL environments and benchmarks to test agent performance on realistic software engineering tasks. Design tasks, review model outputs, and establish success criteria to improve future AI models.
- Remote
- Argentina, Australia, Brazil, Canada, Denmark, Germany, India, Ireland, Netherlands, New Zealand, Norway, Philippines, Poland, Romania, Sweden, Ukraine, United Kingdom, United States
- Posted Aug 20, 2026
- Apply by Sep 19, 2026
- 1 position
More jobs you can apply to directly
Similar opportunities posted by employers hiring on Jobs.ca, with no external application form.
Job summary
About The Job About Surge AI Surge AI partners with the world's leading AI labs to build the benchmarks and RL environments that train and evaluate frontier models. Most benchmarks cut corners because expert humans are expensive. Our model is the opposite: we pay experts enough that we don't have to. Position: Senior Software Engineer (Agentic Coding RL Environments) Type: Contract Compensation: $100–150+/hour ($200–300k+/year). Rates scale with your experience, and many go past the top of that range. Location: Fully remote, any time zone Role Responsibilities Build coding RL environments and benchmarks that test how agents perform on realistic software engineering tasks Design tasks based on real engineering work Review and red-team model output to surface meaningful failures Write the tests, rubrics, and success criteria that separate good work from slop Give structured feedback that improves future models Qualifications Must-Have: Professional software engineering experience shipping real, scaled production systems Expertise in at least one common tech stack (TypeScript, Python, Ruby, PHP, Java, Rust, Go, or C++) Strong judgment about what "good code" and a "meaningful model failure" look like Comfort working independently, and giving direct, code-review-style feedback Availability for a meaningful commitment (most contributors work 20 to 40+ hours per week) Note: You do not need prior AI or ML experience. Strong engineering judgment and attention to detail matter most. What Makes This Different No managers, 1:1s, OKR ceremonies, or Jira backlog grooming You choose when and where you work You are judged on the quality of what you produce, not speed Application Process Coding capability screen (replaces a multi-round interview loop) ID verification Background check Paid first real task. Do well, and you're on the team. Questions? [email protected] Note: Currently, we are seeking Software Engineers from the following locations, and may expand in the near future: Argentina, Australia, Brazil, Canada, Denmark, Germany, India, Ireland, Netherlands, New Zealand, Norway, Philippines, Poland, Romania, Sweden, Ukraine, United Kingdom, and the United States.
What you’ll do
Build coding RL environments and benchmarks to test agent performance on realistic software engineering tasks. Design tasks, review model outputs, and establish success criteria to improve future AI models.
Requirements
Requires professional experience shipping scaled production systems and expertise in at least one major tech stack. Candidates must possess strong engineering judgment and the ability to work independently.
Listed skills
- GoPreferred
- PHPPreferred
- RustPreferred
- TypeScriptPreferred
- JavaPreferred
- RubyPreferred
- C++Preferred
- PythonPreferred
Other relevant skills
Identified from the job description. Confirm important requirements above.
- Software Engineering
- TypeScript
- Python
- Ruby
- PHP
- Java
- Rust
- Go
- C++
- Code Review
- RL Environments
- Benchmarking
- Red-teaming
- Test Writing
Job areas
- Software
- Technology
- Engineering
- Science & Research
- Data & Analytics
Additional details
- Minimum experience
- 5+ years
- Apply by
- Sep 19, 2026
- Posting language
- English
- Working hours
- 40 hours per week
- Location requirements
- Country, Waterloo, Ontario, Canada
- Seniority
- Mid-Senior level
