I
InPost
AI Engineer

Technical Lead, Senior AI Engineer - VonHalsky (m/f/n)

On-siteSeniorAI Engineerposted 4d ago
Role summaryAI-generated

This role focuses on building and scaling an LLM-judge pipeline to evaluate conversational AI performance in production, ensuring real-world Polish-language conversational data is accurately assessed for release decisions. You’ll lead the development of diagnostic tools for root-cause analysis and maintain a high-coverage golden dataset to drive continuous improvement in Von Halsky’s conversational assistant.

Skills required

About this role

Why this role exists

Von Halsky is InPost's conversational AI shopping assistant, live in production and serving a growing share of our customers. What decides whether it wins is not the model but whether we can tell, at release cadence, that a change made conversations better. That is the Evaluations Platform, and we are hiring the engineer who takes it to the next level.

What you will own

  • The LLM-judge pipeline. Our release-gating judge over real and golden conversations, calibrated well enough that people act on its verdicts instead of arguing about it.
  • Eval datasets and the golden set. Real coverage across intents and categories, including the Polish-language coverage generic benchmarks do not give us.
  • Root-cause analysis on real conversations. Making our conversation-mining stack diagnostic rather than descriptive, on a stable issue taxonomy.
  • The "sus" detector. Abusive, adversarial and anomalous sessions, next to our guardrails and red-team work.
  • Evals-driven development. The eval comes before the feature, and writing it is as cheap as writing the code.
  • The interface to product and business. Vague asks in, measurable quality definitions out, and results stakeholders can act on.

What "leadership aspirations" means here concretely

This is a technical lead role, not a people-management role, and you will not carry line-management duties on day one. You will set and defend the technical direction for the platform, act as reviewer of record for the area, mentor other engineers, scope work with our PM and EM, present results to stakeholders, and hold the line against ad-hoc requests crowding out platform work.

Engineering management later, or a Staff-level hands-on track, are both paths we will build with you. Either way we need someone accountable for an area rather than for a ticket.

How we define success in this role

  • Judge pass rate is calibrated against human labels and is a number people trust and cite.
  • A stable, versioned issue taxonomy is live and week-over-week trends are comparable.
  • Every release is gated by an eval run the team can reproduce.
  • The golden set has documented coverage and named blind spots.
  • At least two engineers besides you can operate and extend the platform.

Qualifications

What we are looking for

Required

  • 5+ years building and running production software, with 2+ years on LLM-based systems that real users hit.
  • Strong engineering fundamentals, plus the habits that go with production ownership: testing, CI/CD, containers, observability, and working in cloud. We work primarily in Python.
  • Demonstrable experience evaluating generative systems, not only building them: LLM-as-judge, human-label calibration, inter-annotator agreement, regression suites, offline-versus-online divergence. You should have opinions about what makes an eval worthless.
  • AI engineering fundamentals. Prompt and context engineering as a discipline (context-window budgeting, structured outputs, failure-mode taxonomies); agentic primitives in production (tool use, multi-turn state, MCP, agent-to-agent integration patterns); and eval and LLM-observability tooling (LangFuse, Braintrust, Weave or equivalent, including things you built yourself).
  • Comfort with data at scale: SQL, working with a lake or warehouse, and building a metric someone else can reproduce.
  • Fluency with AI-assisted development tooling (Claude Code, Cursor, Copilot). We use it daily and expect it.
  • Ability to make a technical argument to a non-technical audience and be understood.
  • English B2 and Polish. Our users converse in Polish and you will read their conversations; judging quality you cannot read is not possible.

Nice to have

  • Harness and loop engineering. Building the scaffolding around models rather than only calling them: agent loops, retries and fallbacks, tool-call orchestration, deterministic replay, and the plumbing that makes a non-deterministic system testable.
  • Auto-improving systems. Closing the loop from production signal back into the product: mining failures into cases, using eval results to drive prompt, retrieval and routing changes, and automating the parts of that cycle that people do by hand today.
  • Adversarial robustness, jailbreak testing, red-teaming, or abuse and fraud detection.
  • E-commerce, search or recommendation domain experience.

Additional Information

What we offer

  • A product that has already been released to millions of users, with a real business case, not a lab pilot.
  • Direct access to frontier models at committed capacity across multiple providers, plus an open-source track we run ourselves.
  • A quality mandate with executive attention.
  • Ownership of a platform that is greenfield in practice inside a company with production traffic, which is the rarest combination in this market.
  • Hybrid working from Warsaw or Kraków, in a team growing fast enough that early hires shape how it works.
  • Fulfilling careers with a range of benefits for people and investing in providing training opportunities for their development.
  • You will feel a part of the InPost community that makes an impact on sustainability, convenient deliveries, and the circular economy every day.
  • Excellent working environment and flexible hours
  • We offer B2B type of contract
Apply on InPostOpens in new tab

Similar open roles

T
NEW

AI Engineer, Evals & Agent Quality

Town·San Francisco
On-siteSeniorAI Engineer
$250k – $300k USD
2d ago
TR
NEW

AI/ML Engineering Manager

Thomson Reuters·India, Bengaluru, Karnataka
On-siteSeniorAI Engineer
2d ago
N
NEW

AI Engineer, Agent Platform

NewsBreak·Mountain View, California, United States
On-siteSeniorAI Engineer
$120k – $220k USD
2d ago
CL

Senior AI Engineer

Charger Logistics Inc·Brampton, Ontario, Canada
On-siteSeniorAI Engineer
3d ago
CO

Lead AI Engineer (Gen AI Platform Services: Agentic AI, Agent Guardrails, Agent Evaluation, Agent Memory)

Capital One·San Francisco, CA; McLean, VA; Cambridge, MA; San Jose, CA; New York, NY
On-siteStaffAI Engineer
$197k – $225k USD
3d ago
CO

Staff AI Engineer role (Remote Eligible)

Capital One·San Francisco, CA; McLean, VA; Cambridge, MA; San Jose, CA; Richmond, VA; New York, NY
HybridStaffAI Engineer
$269k – $307k USD
3d ago