Software Engineer 3
MongoDB
- Location
- United States
- Work model
- Remote
- Level
- Mid
- Salary
- $109k/yr
- H-1B history
- 30 approvals (FY2023)
- Posted
- 2h ago
Skills
About this role
The Agent Research and Tooling team, part of MongoDB's AI Builder Experience organization, owns the platform layer around agents: how teams author, distribute, evaluate, monitor, and improve agent skills and agent behavior. We are hiring a software engineer to build and maintain the tooling, evaluation systems, and quality gates behind MongoDB's agent skills.
This is a software engineering role at the intersection of developer tooling, applied AI, and software quality. You will take loosely defined agent and tooling problems, break them into workable plans, and ship durable internal systems: command-line tools, reusable libraries, evaluation harnesses, and CI workflows.
This role is open to remote work in the US or can be based out of any of our US offices.
What you'll do
• Build and maintain agent skills and the infrastructure to validate, evaluate, publish, and maintain them
• Design evaluation datasets and workflows that compare agent behavior against a baseline and produce actionable quality signals
• Build agent metrics and observability: skill selection and routing, success and failure outcomes, tool calls, latency, and token usage
• Design safety and quality gates for agent-authored content: rule packs, static analysis, confidence thresholds, structured verdicts, and bounded suppression
• Create CLIs, libraries, and MCP integrations that other repositories adopt and that run in local development and CI
• Integrate tooling into GitHub Actions and other CI workflows, including secrets, annotations, exit codes, and artifacts
• Build code-generation quality checks, such as anti-pattern catalogs and linting for AI-generated MongoDB code
• Investigate real failures such as nondeterministic results, false positives, and unsafe generated guidance, and turn them into reusable improvements
• Collaborate with engineers, security partners, and product teams; communicate trade-offs, risks, and ownership across teams
Examples of the problems you'll solve
• How can tests verify an agent tool's structured result when item order may vary, but counts, required fields, and values must remain correct
• How can a CI gate flag unsafe instructions in an agent skill without treating every neutral mention as an incident or letting cautionary wording hide a real instruction
• How can an evaluation suite show whether a skill improves answers over a baseline and give authors enough signal to improve it
What we're looking for
• 2+ years of experience building production software, developer tools, internal platforms, or automation systems
• Software engineering fundamentals in API design, testing, error handling, and maintainability
• Experience building CLIs, libraries, test infrastructure, static analysis, or CI/CD workflows
• Ability to design systems that are usable by developers and reliable in automation
• Experience reasoning about correctness and safety with ambiguous input, nondeterministic output, false positives, or untrusted content
• Comfort in an evolving R&D environment where the right abstraction emerges through prototypes and feedback
• Written and verbal communication, including explaining technical trade-offs and aligning stakeholders across teams
Nice to have
• Experience with agentic systems, LLM applications, prompt or rubric-based evaluation, or AI-assisted development
• Experience building eval harnesses, benchmark datasets, quality metrics, LLM-as-judge workflows, or human-review tooling
• Experience with Go, Python, JavaScript/TypeScript, Java, or C#
• Experience with GitHub Actions security, secret handling, static rule engines, or policy enforcement
• Experience moving prototypes into production
What success looks like
In your first year,