Contents · 9 sections
The Role
01 — AI Engineer
Agentic Finance · UK, Luxembourg, Portugal, Spain, Remote-EU · Several roles
AI Engineer, Agent Systems
Noumenai builds AI that does real finance work. Not summarisation. Not a copilot that drafts something a person then rewrites. Agents that go into a financial system, take a position on what a transaction is and — when the evidence and the confidence justify it — post it. Ledger, reconciliations, close, tax. The work that was always done by people because the cost of being wrong was too high to automate.
That cost hasn't gone away. The models are what finally got good enough to make it worth engineering around it. That engineering — not the model — is the company.
We're hiring several engineers into this team, and we're open at more than one level. If you're strong but unsure whether you're "senior enough", tell us where you think you sit and let us judge. We'd rather have that conversation than lose you to a job title.
We're hiring several. Open at more than one level.
The Role
02 — What you'd do
You own a surface, end to end.
Writing code is no longer the scarce resource. Judgment is.
You'll own a real surface of the system — an agent tool family, a connector boundary, an evaluation harness — and you'll own it end to end: design, build, prove it against a live system, and keep it working when reality disagrees with the plan.
You'll work inside an architecture that already exists: the gates, the confidence policy, the agency budget, the tenancy model. You're not asked to invent those from nothing — you're asked to understand why they're there, build well within them, and say so clearly when something doesn't hold.
You'll drive an agent to build most of it. That's how this company works, and it's not optional.
In Production
03 — We are not an idea
There is product in production.
Today, in a client's real accounting, with a live write path into an ERP — reviewer, confidence gates, audit trail. Built in a matter of months.
There are enterprise clients in the pipeline, and the incubation of a joint venture with a multinational group.
The Domain
04 — What this domain actually demands
Agentic finance is not a normal software problem.
These four things apply at every level here, including yours. They are not seniority markers — they are how we work.
01
The failure mode is being confidently wrong.
A model produces a plausible number for anything. Accounting is full of numbers that look right. Our discipline is that the model never originates a value — it chooses the method and explains itself, and deterministic code produces the number. You need to feel why that boundary exists, not treat it as a rule someone handed you.
02
Some actions can't be undone.
A read is free. A posting to a client's ledger is not — some systems don't even offer reversal. Our agency budget tracks reversibility: reversible work runs autonomously, irreversible work is gated at the exit, always, with a reviewer and a human. You'll move fast and you'll be stopped hard, and both are the job.
03
The state of the code is not the real state.
Migration files are not the database. Mocks are not the API. Green CI has meant nothing, repeatedly. Bugs concentrate exactly at the boundaries where someone trusted the shape of a system instead of going to look. The engineer we want is the one who goes and looks — by reflex, unprompted, before claiming anything at all.
04
The domain has opinions.
Accountants and controllers have been doing this for a long time and are right more often than software assumes. When the system disagrees with an expert, the presumption is that the system is wrong. Engineering that can't hold that posture doesn't survive contact with a client.
The Work
05 — What you'll work on
Representative, not exhaustive.
Agent tools and tool families
Designing and building the tools an agent calls: their contracts, their failure modes, their observability. Tool design is product design — a badly shaped tool produces a badly behaved agent.
Connector boundaries
ERPs and financial platforms are hostile, badly documented and idiosyncratic. You'll probe them in a real environment before declaring a boundary functional, and build fixtures from captured real responses, never from imagination.
Evaluation harnesses
Making regressions visible instead of mysterious. If we can't measure whether a change made the system better or worse, we're guessing — and guessing is not available to us.
Features on the write path
Building within the gates: the audit trail, the identity of who acts, the failure semantics. You won't design the trust boundary, but you will build things that depend on it holding.
Data and retrieval surfaces
Institutional memory, grounding, and the retrieval that makes a proposal defensible rather than plausible.
The Fit
06 — What we're looking for
The same four qualities, at every level.
Scepticism, applied to your own work
The rarest quality and the one that counts most. You should feel uncomfortable saying something works before you've watched it work.
Judgment at the boundaries
Knowing which decisions are cheap to reverse and which ones you only get one attempt at — and spending your caution accordingly.
Fluency driving agents
If you're still typing everything by hand, you'll be slow. If you ship whatever comes out of the model, you'll be dangerous.
Honesty under commercial pressure
There will be moments when claiming a capability we haven't built would help close a deal. We don't — not even between ourselves.
Background
07 — You probably already have
What tends to be true of the right person.
- Production experience where being wrong had consequences — payments, fintech, healthcare, infrastructure, trading. Not demos.
- TypeScript and/or Python, and real comfort in a codebase you didn't write.
- Solid Postgres. You should not be frightened by row-level security, even if you haven't implemented it.
More
- Cloud fluency. We're on AWS: ECS/Fargate, SQS, Secrets Manager, Bedrock.
- Claude Code as a daily working instrument — not experimenting, but shipping with it: discovery before implementation, explicit stop gates, one logical unit per PR, diff review before merge.
- Familiarity with the Claude Agent SDK / API — agent loops, tool use, context management across hops.
- MCP (Model Context Protocol) — ideally building servers and tools, not just consuming them.
- LLM systems in production — not prompt engineering, not a RAG demo.
- The habit of reading the source, the docs, or the actual response before claiming something works.
- A quantitative foundation you can reason from, and the ability to read a paper critically when the problem calls for it.
Not Required
08 — What you don't need to have
Four things we don't ask for.
Knowing finance
We'll teach you the parts that matter and shield you from the rest.
A PhD
Not even a Computer Science degree — as long as the foundation is there.
Being a '10x engineer'
A 1x engineer with excellent judgment and a very good model is worth more, and that's what we prefer.
Having architected a trust boundary before
That's what the senior people here are for. You need to understand why it exists and build honestly within it.
The Process
09 — How this moves forward
Two conversations. No whiteboard.
No whiteboard, no algorithm puzzles, no unpaid weekend project.
In the first, we put a real problem on the table — not a sanitised version — and we think about it together, out loud. It doesn't matter whether you reach the right answer. It matters how you think when you don't know.
In the second, you take us to the bottom of a system you built, down to the layer where even you aren't certain any more.
If at any point you tell us I don't know, I'd have to go and look — that counts in your favour.
The conversation you'll want to have
You'll want to know what this actually is — compensation, level, how the day-to-day works, who we are to you and who you'd be to us. We're not going to answer that on a recruitment page. We'll answer it in the first conversation, plainly and without vagueness. What we can tell you now: we're not asking you to trade salary for a promise.
Apply.
Tell us about a system you built where being wrong had consequences — and what you'd do differently now.