Lead Software Engineer - AI/ML - AI Agent Platform
Lead the design and implementation of JPMorgan Chase’s AI Agent Platform, building SDKs, orchestration frameworks, and reusable components that enable conversational, voice, and document-extraction agents across the firm. Drive production quality, infrastructure reliability, and continuous improvement of capabilities such as NL2SQL and RAG for thousands of enterprise users.
Skills required
About this role
Help shape how teams across the firm build AI agents. You will lead engineering for a flagship, high-visibility AI Agent Platform—the SDK, orchestration frameworks, and reusable components that power agentic AI experiences across the organization. Beyond any single chat experience, you will own the foundational tooling that lets teams compose agents—conversational, voice, and document-extraction—ground them in enterprise data, give them durable agentic memory, and continuously improve accuracy on hard capabilities like NL2SQL and RAG. You will raise delivery standards, turn complex problems into reliable production systems, and set the quality bar for infrastructure that many teams and thousands of users depend on. Join a team where strong engineering, thoughtful collaboration, and continuous learning are core to how we operate.
As Lead Software Engineer – AI Agent Platform at JPMorganChase within Enterprise Technology, AI and Machine Learning & Data Platforms, you will lead the technical design and delivery of the SDK and platform capabilities that enable large-language-model-powered agents at scale. You will own the building blocks agent orchestration, tool/function calling, retrieval and grounding, voice interaction, document extraction, and agentic memory and drive measurable, iterative improvement ("hill-climbing") of capabilities such as NL2SQL and RAG. You will translate high-impact problems into production-grade solutions, from discovery and design through deployment and ongoing operations, while partnering closely with product and stakeholder groups to deliver measurable outcomes.
Job responsibilities
- Lead end-to-end delivery of the AI Agent Platform SDK and frameworks—from problem framing and technical design through production deployment, scaling, and monitoring—that engineering teams use to build, evaluate, and operate AI agents.
- Own core platform services and reusable building blocks: agent orchestration, tool/function calling, retrieval-augmented generation pipelines, and model-serving integration.
- Build and operate first-class agent modalities on the platform, including voice agents (speech-to-text, text-to-speech, low-latency streaming, and turn-taking), document-extraction agents (parsing, OCR, structured field and table extraction from complex documents), and agentic memory (short- and long-term memory, persistence, retrieval, and context management across sessions).
- Drive systematic hill-climbing of core agent capabilities—including NL2SQL, RAG, voice, document extraction, and tool use—by building evaluation datasets, benchmarks, and quality metrics, then iterating on prompts, retrieval, and orchestration to measurably improve accuracy.
- Establish engineering standards through hands-on system design, rigorous code review, and mentorship to improve reliability, maintainability, latency, and developer experience for teams building on the SDK.
- Build and operationalize evaluation, testing, and observability capabilities (offline eval harnesses, regression suites, tracing, latency and cost telemetry, logs, and analytics) to continuously improve solution quality.
- Implement robust safety and governance patterns—guardrails, content filtering, prompt-injection defenses, access controls, and audit-ready operational practices aligned to enterprise expectations.
- Partner with product managers and stakeholders to shape the roadmap, define success metrics (capability accuracy, adoption, task completion, quality), and prioritize work that delivers measurable business impact.
- Drive cross-functional alignment across engineering, data, security, and risk partners to ensure the platform is secure, stable, performant, and scalable.
- Contribute to technical documentation, SDK references, reference implementations, and enablement content that accelerates adoption and responsible usage.
Required qualifications, capabilities and skills
- Formal training or certification on applied artificial intelligence and machine learning concepts and 5+ years applied experience.
- Advanced proficiency in Python with strong software engineering fundamentals, including testing, design patterns, version control, and code review practices.
- Hands-on experience building, evaluating, and deploying machine learning or large-language-model-enabled systems into production environments—ideally reusable libraries, SDKs, or platform services consumed by other engineering teams.
- Practical experience with prompt engineering and retrieval-augmented generation, including building evaluation methods, benchmarks, and quality measurement to systematically improve capabilities such as NL2SQL and RAG.
- Experience designing and operating reliable services, including incident response readiness, performance tuning, and operational stability for data-intensive systems.
- Demonstrated ability to lead technical decisions and deliver outcomes through ambiguity, balancing speed, risk, and long-term maintainability.
- Strong communication skills with the ability to explain technical trade-offs to both technical and non-technical stakeholders.
Preferred qualifications, capabilities and skills
- Experience with agent orchestration frameworks (for example, LangGraph, LlamaIndex, Google ADK, or custom orchestration) and evaluation tooling for LLM systems.
- Experience building voice agents (speech recognition, text-to-speech, real-time streaming audio, and conversational turn-taking).
- Experience building document-extraction / intelligent document-processing agents (OCR, layout parsing, structured field and table extraction from complex or unstructured documents).
- Experience implementing agentic memory systems—short- and long-term memory, persistence, and context retrieval across sessions.
- Familiarity with vector databases, embedding pipelines, or graph-based memory approaches used in retrieval-augmented generation solutions.
- Experience with continuous integration and continuous delivery practices and containerization (Docker and Kubernetes) for production deployments.
- Experience with cloud and machine learning platforms (for example, Amazon Web Services, Databricks, or comparable platforms).
FEDERAL DEPOSIT INSURANCE ACT: This position is subject to Section 19 of the Federal Deposit Insurance Act. As such, an employment offer for this position is contingent on JPMorganChase’s review of criminal conviction history, including pretrial diversions or program entries.
#LI-RB3
#AMDPAIML