Systems · Open role


Build the runtime that makes agent output trustworthy.

A confident-sounding model isn't a system a reviewer can trust. The runtime is the difference: the extraction pipeline that ties every value to its source, the validators that run before output lands, the evals that stop a bad build. You build it. The stack is Next.js and Convex.

Mission


What the role does and why it matters here.

You build the system that makes agent output worth a reviewer's trust. The extraction pipeline ties every value to its source page. The evals catch a build that shouldn't ship. The validators the Context Architect defines live in code, run on every step, and feed the audit trail. It has to work the same way across engagements, on imperfect source material, at the standards regulated enterprises run on.

Responsibilities


What you would own.

Build and maintain the agent runtime

The harness brings together an agent's tools, memory, and retrieval. You build it so the same runtime works across engagements with different validator sets. It lives in the IAS codebase and in engagement-specific repos.

Own the extraction pipeline

Every value the agent extracts carries a confidence score and a link back to its source page, so a reviewer moves from value to source in one workflow. You build and maintain the pipeline behind that, including how it behaves on inconsistent exports, scans, missing fields, and documents that disagree.

Run the eval suite

Truth sets the team agrees on. Regression coverage on changes. Drift alerts when the harness behaves differently from a prior run. The bar is an eval that has actually stopped a build that should not have shipped; you build toward that bar and away from evals that pass vacuously.

Ship validators as code

Working from the Context Architect's spec, you implement the validator set as code that runs alongside every agent step. Failures are visible. Advisory checks are marked as advisory. Validator output is part of the audit trail.

Maintain the provider abstraction

Mock providers during development, real providers in production, no rewrite in between. The interface is yours to design; the migration path is part of the design.

Build production observability for agent runs

The team has to be able to read what an agent did before they trust the next run. You build the observability that supports that, without surfacing harness internals in the user-facing UI.

How you think and work


Six traits the work demands.

Pedigree isn't the filter. Disposition is. The six traits below are what the work actually asks of you.

  1. Agentic intuition

    You read agents the way a manager reads a direct report: when to trust the output, when to interrupt the run, when to take the wheel back.

    The evals you've shipped are the ones that have actually caught a run that should not have shipped, not evals that pass because they checked the wrong thing.

  2. Critical thinking

    Confident-sounding output gets the same scrutiny as anything else, your own work included.

    You've found an eval that was passing for the wrong reason, and rewrote it.

  3. Curiosity

    You pull on threads. You read outside the lane. You follow a question past the first plausible answer.

    You'd rather read the implementation than the README.

  4. Agency

    You move without being told. You decide, ship, own the call. No one has to write the playbook for you.

    You've shipped the change because the meeting would have taken three weeks.

  5. Systems thinking, long view

    You see how the parts connect, and where this goes in three years.

    You build the interface knowing the second use case will teach you more than the first.

  6. Leadership instinct

    You orchestrate work across humans, agents, and stakeholders. You switch register between a workspace ticket, an architect call, and a senior bank room in the same day without losing what you came in to say.

    You've explained a tradeoff to a non-engineer and walked out with the right call.

Useful background


  • You've built a regression discipline (eval suite, golden tests, drift monitoring) that has actually caught a release that would have shipped wrong.
  • You've handled messy real-world inputs: inconsistent source documents, scans, dirty cross-system data, or records that disagree. You know where it broke and what you did about it.
  • You've shipped a typed end-to-end stack to real users. The Next.js + Convex stack we use is one shape of this; we read the discipline, not the framework.
  • Your prompts, or the inputs to your reasoning systems, live in version control with diffs and review. Because the alternative bothers you.

Regulated-industry experience isn't required. Curiosity about it is.

Logistics


How the role is set up.

Engagement type
Contract or full-time. Contractors run on a defined engagement scope and can convert to full-time. Full-time runs on a yearly review with a quarterly written check-in.
Location
Remote. Lucentive is EU-based; we expect at least four hours of overlap with CET on a working day. Travel for engagement kickoffs is occasional, not weekly.
Compensation
Discussed in the written exchange, at market rate for senior engineers shipping AI into production. The band depends on contract or full-time and on location.

Apply


We look forward to hearing from you.

Apply for AI Engineer