Remote job
Remote confirmedSenior Applied AI Engineer, AI Platform
Bunch
Key points from the posting
- Tech stack
- TypeScriptSvelteReactNode.jsNest.jsMySQLPostgreSQLMastraVercel ai-sdkGoogle Vertex (Gemini)AWS BedrockKubernetes
- Seniority:
- Senior
Read out of the job posting automatically
Our assessment
- Our reading of the full posting text confirms it: fully remote.
- 19 more open roles from this employer in our index. 1 of them fully remote.
This section only: calculated automatically by nomado24, from our own job index and our own reading of the posting text. Not stated by the employer.
Job description
bunch https://www.bunch.capital/ is building the backbone of private markets. We are enabling next-gen fund operations with one integrated system that combines secure data infrastructure, AI-powered workflows and expert fund services. If you value ownership, growth through real responsibility, and working with a thoughtful, ambitious team, this role might be for you.
YOUR ROLE
As a Senior Applied AI Engineer on our AI Platform team, you take AI at bunch from working to relied upon. We already run AI in production — a document extraction pipeline live in fund operations, agent workflows on Mastra, evaluations, and multi-provider fallback inside EU data residency. Fund operations run on documents, deadlines and numbers that have to be right — subscription documents to parse, capital calls to chase, portfolio data to reconcile. You build on that foundation: more agents taking that work off people's hands, the evaluations that prove they can be trusted with it, and the platform that lets every other team at bunch ship the same way. This is end-to-end product engineering, not research: you own architecture, evaluation, integration, and deployment.
TOP PRIORITIES
- Build and ship agents. Design agents that automate real fund-operations workflows, and own them from prototype through production and after. They integrate with our services, data model, and authorization system — they don't sit beside the product as standalone prototypes.
- Evaluate and improve agent performance. Build the evaluation layer: test cases built from real documents with the output we expect, regression suites in CI, human review where correctness is non-negotiable, and clear success criteria for an agent completing a complex task end to end. Then move the numbers that matter — accuracy, latency, cost.
- Own the AI application architecture. Orchestration and multi-agent design, tool contracts, memory, and context engineering (RAG, MCP) with clear domain boundaries — plus the guardrails, approvals, and human-in-the-loop controls that anything touching investor money requires.
- Run it in production. Versioned, feature-flagged rollout of agent versions; rate limits, provider fallback, and regional failover within EU data residency; and the observability to trace a failure across services and turn it into a fix rather than a theory.
- Make it a platform, not a project. Document parsing and extraction consolidates into this team, and shared evaluation and observability become something other teams consume rather than rebuild. You set the patterns other engineers inherit, partner closely with DevX, and mentor engineers across teams on agent and LLM practice.
YOUR FIRST 90 DAYS
- Review our MCP offering end to end and ship it to the first customers, with the access boundaries, evaluation and observability a customer-facing surface needs.
- Stand up shared observability for our AI workloads — token usage, estimated cost, latency, failures and retries per provider — on a dashboard people actually open during an incident.
- Publish v1 of our agent patterns (domain boundaries, tool contracts, prompt and eval conventions), reviewed with DevX and adopted by at least one team outside AI Platform.
ABOUT YOU
- Experience: 5+ years building production software, including at least one agent or LLM-powered capability you took end to end and still owned once it was live
- Agents: real depth in orchestration and context engineering — tool contracts, memory, RAG, multi-agent design — with Mastra, ai-sdk, LangGraph or similar. You've integrated agents into a real product and its authorization model, not built prototypes for someone else to productionise
- Evaluation: you know how to make a non-deterministic system measurable, from test cases built on real data to regression suites and human review, and you can tell the difference between a model that improved and a benchmark that got easier
- Production: you own what you ship. Reading traces, diagnosing failure modes, tuning cost and latency, handling provider rate limits and fallback without drama
- Backend: You have experience with TypeScript and/or Python.
- Platform mindset: you build for other engineers as much as for end users, and the standards you set get adopted because people trust you, not because they are written down
- Pragmatic: you start from the business outcome, choose the deterministic solution when it's the right one, and know when good enough is good enough
- Experience in fintech, private markets, or another regulated, document-heavy domain is a plus, as is working under EU data-residency constraints
OUR TECH STACK
- Frontend: TypeScript, Svelte and React
- Backend: Node.js, Nest.js
- Database: MySQL, PostgreSQL
- AI: Mastra, Vercel ai-sdk, Google Vertex (Gemini) with AWS Bedrock fallback in EU regions
- Infrastructure: Kubernetes on AWS
- Observability: Datadog
- Auth & Internal tools: FusionAuth, Retool …
This role is provided by an external source. Applications are handled on the source website.
