Forward Deployed Engineers · 6 min read
We Spoke to 100 Forward Deployed Engineers. This Is What We Learned.
What a hundred interviews taught me about the difference between building an AI system and running one, and the one question that shows which of the two you are hiring.
Bram van Gestel · Published 2026-10-07

My daughter is eleven. A while ago she asked Gemini for a gluten-free cookie recipe, swapped in what we had in the cupboard, and baked twelve. They were good. I mention it because that is where the tools are now: anyone can get a model to produce something that works once.
It is also the problem with hiring AI engineers this year. Most CVs say "built AI systems". So did my daughter, in a sense.
I run AgentLabs with my co-founder. We build agentic systems inside other companies, and some clients are best served by one of our engineers working in their offices, on the systems they already run, next to the people whose work is about to change. The market calls that a forward deployed engineer. To find them I have now spoken to a hundred engineers who want the job. This is what stayed with me.
Who used it?
That is my first question, followed by what happened the first time the system did something wrong. An engineer who has shipped answers with a story: the documents that turned out to come in two layouts, the fields the business kept adding after go-live. An engineer who has not shipped answers with architecture.
About fifteen minutes into one interview a candidate stopped halfway through describing a system, a careful one, to tell me it had never gone live. Research, tested on public data, never run by a company. I had not asked. After that I trusted every answer that followed.
The second question
Nearly everyone is fluent for one answer, so I ask a second one and I keep it dull on purpose. How does this webhook behave when the other side times out, and what happens when the same message arrives twice? If I ask you tomorrow what the agent did on Tuesday, which sources it read and what it cost, where do you look?
These are the questions a client's IT manager asks me, so I ask them first.
The answers thin out in the same places: retries and signature checks on an integration, and tracing what an agent did and what the tokens cost, with Langfuse or with a log you built yourself. Deploying into a company's own stack was the other weak spot. Nobody goes vague about the models. Coding is solved; the layer above it is not.
Cost is where I have the least patience, because I have watched companies put an LLM on PDF parsing that a Python package would have done for free, and find out from the invoice. An engineer who has been through that once brings it up without being asked. It came up less often than I would like.
Twenty years of finding my own mistakes
At my last company I moved a team of forty engineers to a way of working where they no longer write code by hand. What they build instead is the harness: the gates, the validators, the review that fails. The engineers who were best at it had maintained software long enough to know where it rots.
So when a candidate tells me how they keep a coding agent under control, I listen for things I can go and look at: warnings turned into build errors so the agent cannot ship a shortcut unnoticed, a validator that fails the build when one module reaches into another, a review step that scores each change and hands anything under the threshold to a person.
I met the opposite too, more than once. Engineers a few years in, often with a machine-learning degree, fluent in Claude Code or Codex, who trust the model and cannot tell me how they would measure what it produced. I understand how that happens. If you have never had to find your own mistake in a running system at a client, the model's confidence is very persuasive. I am still not looking for vibe coders. A client's systems are old, its data sits in several places, and somebody has to know where software goes wrong before it goes wrong there. Architecture, software systems and infrastructure come first, and AI experience on top of that.
What I trust
One candidate told me plainly they had never used tooling to trace what an agent did. Another described letting a coding agent run unchecked on a side project until every boundary in the code had turned into untyped strings and the whole thing had to be redone. On a checklist both lose points. Both are people I would put in a client's building, because they will tell me in week two that something is not working.
The best question anyone asked me was whether you can spend months on a project and end with nothing that works. Yes. The usual reason is data in worse shape than anyone admitted at the start, and I have assumed too much about a client's data myself.
The people whose work it was
The last few years almost every organisation tried some kind of AI transformation. Many did a decent job on the systems and forgot the people. The consultants left, and the people whose work it was looked at each other, and then went back to doing it by hand with an agent somewhere doing nothing.
So the second interview, with my co-founder, is one scenario: a person who does a task by hand for two hours every morning and has every reason to distrust the new thing. I was wrong about who does well there. I assumed the founders among the candidates would be strongest. The best evidence came from a machine-learning engineer who ran weekly sessions with the team that would use the system and turned the sceptics by showing it working, and from a researcher who had spent years explaining models to specialists whose requirements kept moving.
Why they want it
Some wanted colleagues again after building alone. Some were at the end of a contract. Some wanted more agent work and less platform work. One told me their energy drops once a job settles, and I started this company because mine does too: after twelve to eighteen months somewhere, if I have done my job, it turns into shopkeeping.
What they share is the wish to own something from the first conversation to the day it is used. One candidate put it better than any job description I have read: it is a job for people who want to own a task from start to finish.
If you are hiring one yourself
Three questions do most of the work. Who used it, and what happened the first time it failed? How did you know it was right? What would you check before letting it act on a live system? Ask each one twice.
The engineers we selected this way are working at clients now. We hand-pick every one of them, and we do not leave them alone on site: regular check-ins on process, progress, what the agents can and cannot yet do, and coaching. When the work is done the client's team runs it, and we leave.
If you have a seat like this open, send me a message.