How to Vet Your First Marketing Engineer Hire
The interview loop, the work sample, the questions that reveal who can do the work, and the signals you can read in the first call.
“How to Hire Your First Marketing Engineer” is the most-read piece I have published in the last three months, and it’s crushing for SEO. It went out in April. People still email me about it, forward it to their VP, and lift the JD out of it. (Lift it. That is what it is for.) A piece like that names a gap, and what usually comes back is more posts about the gap. Profound went and built into it instead. They’ve shipped a job board for the role, a certification program that makes candidates build and defend a working agent on camera, and a hackathon where the prize was an interview. The rest of the category is still arguing about whether the title is real.
Which is funny, given what I did not put in that piece.
When I first heard the title going around LI, I rolled my eyes a bit. My read was that somebody in Clay’s orbit had taken GTM engineering, swapped a word, and shipped a category. I have watched vendors invent job titles to sell software my entire career, and I assumed I knew exactly what this was.
I was half right (which is the most expensive way to be right). I did not wave the role off because I did not understand it. I waved it off because I thought it was my resume with a new label on it.
I came up in marketing operations, which meant I spent my early career inside the machine rather than in front of it. Lead routing, scoring models, the whole path from a form fill to a closed-won deal, instrumented well enough that you could see where it leaked and prove it to someone who did not want to believe you. That was always systems work. The people who were good at it were good because they could hold the entire path in their head at once and tell you which part of it was lying to you.
Then I spent the next fifteen years bolting business judgment onto that. Global digital marketing at FIS across a dozen markets. Scaling Affirm through IPO. Building growth at Webflow. Every one of those jobs was the same exercise at a bigger number: understand the system, find the part that actually moves money, go change it, defend the result to a room that has its own theory.
That combination is my whole career advantage and I have never been shy about it. Not the automation, not the strategy, the fact that I have both when most people pick one. I could read the system and I could read the P&L, so I knew which lever was worth pulling, and I could usually go pull it myself instead of filing a ticket and waiting a quarter.
Which is exactly why the new title did not land for me. Systems thinking plus business judgment plus a willingness to go build it yourself is not a new job. It is the job I have been doing since I was in my twenties.
Then the aha, and it was not that I had the role wrong. I had the scope wrong. It is now the role I am most bullish on in all of marketing.
The title is not inventing anything. It is putting a name on what a stellar marketer already is right now, and what the rest of the function is about to have to become. That combination used to be an edge, the thing a handful of us had because we came up inside the machine instead of in front of it. It is not an edge anymore. It is the baseline for anyone who wants to be excellent at this in three years.
The skill did not change. The ceiling did.
Everything I used to need a team and two quarters for, I can now do alone in an afternoon. I have shipped a product this year with auth, payments, quota enforcement at the database layer and its own MCP server, solo. I built a client’s audit tool in a week that would have been a scoped engagement with a kickoff call. The site, the analytics, the email infrastructure, the launch assets. None of that required me to become an engineer. It required the thing I already had, pointed at tools that finally stopped making me ask permission.
So the constraint moved. It used to be capacity: I knew what to build and had to negotiate for the people to build it. Now capacity is close to free and the only scarce input left is knowing what deserves to exist. I can build almost anything I can specify inside a week, and I am wrong about what to build far more often than I am wrong about how to build it. No tool is coming to fix that.
That is what the role actually is, and why it is not GTM engineering with a new coat of paint. The best marketers I know right now are not specialists anymore. They are a small org compressed into one person, running four things at once: the creativity to come up with something worth making, the judgment to know which of those ideas earns the week, the systems thinking to see how the pieces move each other and where the whole thing leaks, and enough tooling fluency to go make it exist by Friday without asking anyone. That used to be four people, a roadmap and a standing meeting. Now it is one seat, and it is the most leveraged seat on a marketing team.
I think that is where the whole function is going. I also think the people who get there first will be the ones who have already spent years feeling the system underneath the tools, because when the tools change again, and they will, the systems thinking is the part that transfers.
None of this is theoretical either. I was talking with Harvey this week and the role came up on its own, unprompted, because they have one open right now. That is the tell. When a company like that decides it needs this seat before there is a playbook for filling it, the demand is already ahead of the operating knowledge.
The search data says the same thing, louder.
“What is a marketing engineer” is a breakout query on Google Trends in the United States over the past month.
Breakout is a defined term in Trends, not a figure of speech. It means the search grew more than 5,000% against the prior period, the label Google uses when a term is new enough that there was almost no earlier volume to measure against. Translated: a wave of people found out this role exists in the last thirty days and went looking for what it means. “Marketing engineer” on its own is still climbing.
People are hiring for this faster than anyone has worked out how to hire for it. I am writing this to close some of that gap.
Which brings me back to the point I opened with. Most of this category still lives in scattered LinkedIn posts and a few Slack threads, and the operating knowledge stays thin because almost nobody funds the unglamorous parts. Profound is the exception, and the job board and the university are the newest pieces rather than the whole of it. They are the analytics and action layer big teams (Ramp, US Bank, Rippling) use to see how they show up inside AI answers and then go change it. Read that against the definition I opened the last piece with. The marketing engineer closes the loop alone: find the problem, build the fix, ship it, measure whether it worked. Measurement plus the means to act on it is that loop with tooling under it, which is a more useful thing to hand someone than a dashboard they can only read.
None of that infrastructure can make the hire for you.
Scoping the role is easy. Writing the req is easy. The hard part is forty-five minutes across a table from someone with a clean resume, a slick portfolio, and a browser tab open to every AI tool you have heard of, where you have to work out whether they can do the job or whether they are good at describing it.
No role in marketing is harder to read right now, and nobody is to blame for it. The work is new and it will not sit still, so the signals every hiring manager spent a career learning stop predicting whether someone can do the job. A standard loop is not built to catch the difference. It rewards the wrong candidate without anyone meaning to.
This piece is how to build a loop that shows you the work instead. At the bottom you get all of it as a kit you can lift: the screener script, the work-sample brief word for word, the full question bank, the scorecard, and the calibration rules for the debrief. Paste it into Greenhouse and run it Monday.
TLDR of the first piece
If you missed it, the argument in six lines.
The role is builder and artist in one seat. They close the loop alone: find the problem, build it, ship it, measure it, kill it or scale it. What used to be a cross-functional project is now one person’s week.
The bar is creative judgment, not technical skill. The tools are easy and getting easier. Knowing what to build is the job.
They report to the CMO or VP of Marketing. Not ops, not a chief of staff, not a dotted line into engineering. Team of one for the first six months.
They own what they build and influence everything else. Give them a channel to own and you have built a demand gen manager with extra steps.
Pay senior IC, not marketing band. $180K to $220K base plus senior IC equity. You are competing with engineers and PMs, not with marketers.
Two failure modes to design against. They ship nothing under ambiguity, or they ship constantly and nobody uses any of it. The second is a taste failure and it hides the longest.
That first piece was about scoping the role and writing the req. This one is everything between the req and the offer.
One resource before we start. If you are still sourcing, Profound runs the only curated board for this role. Worth a look for two reasons. The live roles (Figma, Stripe, Databricks, Plaid, Legora) tell you who is hiring, which is useful competitive intel when you are trying to close someone. And the posted comp bands run roughly $100K to $296K right now, which is a free market read before you write your own offer.
Why this role is so hard to read in an interview
I have made hundreds of hires. Enough that I can usually read a candidate inside ten minutes, and enough to know which signals I am reading while I do it. Paid marketers, content leads, demand gen managers, ops people, PMMs. For every one of those I walked in knowing what competence looked like from across a table, because I had seen it and its opposite often enough to tell them apart fast.
Most marketing work is legible. I could read it quickly for so long because it holds still. Look at a paid marketer’s accounts, spend, and ROAS and you know inside an hour whether they ran serious programs. A content lead hands you a portfolio and you read it. A demand gen manager walks you through a pipeline model and you can tell if the math holds. The signal sits still long enough to read, and after enough reps you read it without thinking.
The marketing engineer’s work does not sit still. The tools are new. The field has a history you can measure in months, so the vocabulary is young enough that fluency and competence sound identical in conversation. A working demo and a production system look the same on a screen.
So every signal I spent a career learning stops predicting the one thing I care about.
Knowing every tool used to be a strong signal. Cursor, n8n, Clay, v0. Now the tools are easy, so tool fluency and judgment sound the same across a table and only one of them ships work people use. A demo that runs beautifully in the room tells you the happy path works, which tells you nothing about a Tuesday in production, and Tuesday in production is where the job lives. Someone can describe an agent in language that sounds like architecture and fully believe they built one, when what they shipped was a single workflow with a model call in it. They are not lying to you. The distance between those two things is invisible in words. And everyone can talk fluently about creative judgment now, because it is what everyone says the role is about, which means saying it and having it have become impossible to separate in a conversation.
The field is not full of pretenders. The field is new, the best work is hard to see, and the interview methods all of us inherited were built for roles where competence announces itself. It does not announce itself here.
The fakers never bothered me as much as this did: the same blindness that lets a talker through will also lose you someone good.
I have been on the wrong side of this, in the other direction. Early in my career I hired for the tool, because the tool was the only thing I knew how to test for. The person could administer the platform faster than anyone else I interviewed. What they could not do was tell me what the platform should be doing, and I did not find out for two quarters, because somebody who cannot find the leverage still looks busy. This cost me months but also taught me a valuable lesson…I started hiring for marketing judgment first, even when tool knowledge fell slightly short.
So the fix is not a more suspicious interview. It is a more revealing one. Stop asking candidates to describe the work, which even the strong ones find hard, and give them a way to show it.
I have seen this movie before, in marketing ops
Marketing automation arrived, the martech stack went from a handful of tools to thousands, and every team suddenly needed someone who could run the machine. Companies did the obvious thing. They hired for the tool. Can you administer Marketo? Do you know Salesforce? Have you run Eloqua? It worked the way hiring for the tool always works. It filled the seat and missed the gap.
People who could click through the platform were everywhere. The ones worth hiring were rare, for reasons that had nothing to do with the platform. They understood the system. They had judgment about what to build and, more importantly, what to leave alone.
Then the tool knowledge commoditized. Within a few years everyone had it and it stopped being worth anything.
The judgment never commoditized, and neither did the systems thinking underneath it. Operators who had both got more valuable as the stack got more tangled, because someone still had to understand how the pieces moved each other before deciding what the machine should do. The understanding is the part that transfers. Somebody who can see how a system fits together can point that at a stack they have never touched, which is why the good ones stayed good through three generations of tooling while the tool experts got replaced with each one.
The marketing engineer is that person with better tools and a higher ceiling. Swap Marketo and Salesforce for models, agents, and pipelines and the shape is identical. What separates a great one from an expensive one has not moved in fifteen years. It is judgment about what to build, and a feel for the system they are building inside of.
One clarification, because the analogy invites a wrong read. The marketing engineer does not replace marketing ops. Marketing ops is still its own discipline, still holding weight, still keeping your lead flow, routing, attribution, and systems of record from falling over. This is not a succession where one role retires and hands over the chair. The same pattern is repeating one layer up, and on a mature team the two need each other.
Which is also why the mistake is identical. Hire the marketing engineer who dazzles you with tools and you have made the error teams made hiring ops in 2012. You will fill the seat and miss the gap.
The one rule: audition, do not interview
Every good filter in this loop comes back to one move. You put the candidate in a position where they have to build something, decide something, or kill something, in front of you, in real time, with a wrong answer available.
Description is where the strong and the not-yet-ready sound alike. The moment the work goes live, they stop sounding alike. The gap between someone who has shipped fifty tools and someone who has mostly read about shipping them is invisible in conversation and impossible to hide in a build.
So the loop is built around three things that are genuinely hard to fake:
A live build under constraint. Not “tell me about something you built.” Build something now, or on a short take-home, and let me watch how you work.
A judgment call with a wrong answer in the room. Build versus buy, with a real tradeoff, where the impressive-sounding answer is the wrong one.
A walkthrough of something they already shipped. Have them demo a real thing they built and drill into the thinking behind it. What did they choose, why, where did they get stuck, how did they get unstuck. The story falls apart fast if they did not do the work, because it only holds up in specifics the builder would know.
Everything below is in service of those three. Skip them and you are back to hiring on vibes.
Profound ran a very interesting version of this. They held a marketing engineer hackathon and made the prize an interview. You build under a clock, in front of people who know the work, and if you win you get a shot at the seat you were auditioning for the whole time. That is the loop, compressed, and it runs before anyone has spent a dollar on a screen. The build already happened. Anything the candidate says in the room afterward is confirmation of something you watched, not a substitute for it. The part worth stealing has nothing to do with hackathons. A job post pulls people who are good at applying for jobs. A build competition pulls people who want to be judged on the build, and that is a smaller and far better pool, sorted for you before you read a single resume. You can get some of this without the event. Put the work sample first, ahead of the screen, and say so in the posting. One way or another, make people show you what they can build in a short window. The ones who opt out were going to cost you four rounds to find out about anyway.
The loop
Four stages that matter, plus the posting and the debrief that bracket them.
Stage 1. The screen. Thirty minutes, run by the VP or a strong IC. Not a resume walk. You are listening for a short list of tells, below, that decide whether to keep going. Most of the field ends here and it should.
Stage 2. The work sample. The audition, and the most important stage in the loop. A small scoped problem, either a 90-minute live build or a three-hour take-home with a hard cap, paid. What matters is not the artifact. It is how they work.
Stage 3. The judgment loop. Two 45-minute rounds. One on what to build and build versus buy. One on measurement and taste. This is where you find out whether the person who built well in Stage 2 knows what is worth building at all.
Stage 4. References. Two calls, with questions that are not the usual ones. Almost nobody does this well.
What I left out: no take-home that eats twenty hours, because that filters for the desperate rather than the good. No brain-teasers. No panel of six people asking overlapping behavioral questions. Every stage earns its place by testing something a resume cannot.
The work sample: how to run the audition
This is the most important forty-five minutes to three hours in the loop, so run it with intent.
Give them a real problem, small and scoped. Not a take-home platform build. A single workflow with a clear edge. Examples that work:
Here is a folder of ten customer call transcripts. Build something that turns them into a usable content brief. You have three hours.
Here is our category and the five prompts a buyer would actually type into an AI tool to research it. Show me how we show up in the answers today, find the gap that is costing us the most, and build something that starts to close it. Tell us what you would measure to know it worked.
Here is a competitor’s pricing page and ours. Build an agent or tool that flags a meaningful difference and drafts a one-paragraph response narrative.
Here is a messy CRM export. Build the enrichment and scoring pass you would actually run, and tell us what you would not trust in the output.
Watch the front of the work, not the end of it. The finished artifact matters least. What tells you everything:
What they clarify before they start. A strong one asks two or three sharp scoping questions and then goes quiet and builds. A weak one either starts flailing immediately or asks for a spec. The instinct to define the problem themselves is the whole role. If they need you to define it, you found a marketing ops manager with newer tools.
What they cut. In three hours, nobody ships the whole thing. The good ones cut deliberately and tell you what they cut and why. The weak ones try to build everything and finish nothing, or polish one corner and leave the core broken.
How they narrate tradeoffs. “I used the cheaper model here because this step does not need the reasoning and it runs a hundred times a day” is the sound of someone who has paid a production bill. “I used the newest model because it is the best” is the sound of someone who has not.
Whether it actually runs. Not “here is what it would do.” Runs. On real input. With a failure mode they can name.
Then have them break their own work. The last fifteen minutes: “Where does this fall over? What breaks it at 2am? What would you not trust?” Someone who has run real systems answers instantly and specifically, because they have watched their own break. Someone who has mostly built demos says it is pretty robust. That answer is the whole interview.
The questions that carry the weight
Group your judgment rounds around five things. For each question below: what it is really testing, what a strong answer sounds like, and the red flag.
Use these live. The follow-ups matter more than the openers, so keep pulling the thread.
On taste (what to build)
This is the load-bearing skill, so it gets the most airtime.
“Walk me through the last three things you built. For each one, tell me what you decided not to build to make room for it.”
Testing: whether they think in tradeoffs, and whether they kill things. Building is a sequence of choices about what not to build.
Strong answer: they name the thing they skipped, and the reason is about leverage, not time. “I did not build the dashboard because nobody was going to open it daily, and a Slack alert did the same job in a tenth of the effort.”
Red flag: they cannot name anything they skipped. Everything they touched, they built. That is not range, it is no filter.
“Show me something you built that nobody used. What did you learn?”
Testing: honesty and volume, together. Everyone who ships a lot has built dead tools. It is unavoidable.
Strong answer: an immediate, specific, slightly embarrassed story about a tool that looked useful and rotted. Bonus if the lesson changed how they scope now.
Red flag: they cannot think of one. That means either they have not shipped enough to have a graveyard, or they are not being honest about the one they have. Both are disqualifying, for different reasons.
On build versus buy (the judgment)
“We are about to renew [real tool] at [real number] a year. Would you build the replacement? Walk me through how you decide.”
Testing: whether they reflexively say “build it” to impress you, or actually reason about maintenance cost, reliability, and opportunity cost. This is the trap question. The impressive-sounding answer is the wrong one.
Strong answer: they do not answer the question. They ask about how load-bearing the tool is, who depends on it, what breaks if the internal version has a bad day, and what else they would not be building if they built this. Then they might say buy. A marketing engineer who is comfortable saying “keep paying for it” is more senior than one who wants to build everything.
Red flag: “Oh, I could build that in a weekend.” Maybe they could. The point is they did not ask a single question about what it would cost to own it for two years. Building is cheap now. Owning is not.
“Tell me about a time you decided to buy instead of build.”
Testing: whether they can kill their own instinct to build. A builder who always builds is a liability, not an asset.
Strong answer: a real story where buying was correct and they knew it, usually because the thing was not a differentiator and the maintenance was not worth the ego.
Red flag: they cannot find one, or the story is really a build story wearing a buy costume.
“What is something on your team that could be automated, that you would leave manual anyway?”
Testing: whether they price automation against what it costs to own. Building is cheap now, so the discipline is not in what they can automate, it is in what they choose not to. This is the same judgment as build versus buy, one level down, at the task instead of the tool.
Strong answer: they name a specific task and the reason is ownership cost, not difficulty. The volume is too low to earn the maintenance, the inputs change every quarter so it would break constantly, or a human catches something in the doing that the output would lose. If they do not have a ready example, a candidate who reasons one out in front of you is answering the question just as well. The thinking is the thing being tested.
Red flag: they cannot find one, in memory or in the moment. Everything is a candidate for automation and nothing costs anything to keep alive. That is the disposition that fills a team’s week with agents nobody maintains.
On killing work and reliability
“An agent you built has been running for four months. How do you know if it is still worth running?”
Testing: whether they think past the ship date. Most of the failure in this role is not bad builds, it is good builds nobody maintained.
Strong answer: they talk about monitoring, drift, unit cost creep, and a specific check they run to decide whether to keep it alive. They treat their own tools as things that can rot.
Red flag: the question surprises them. They ship and move on. In this role, that is how you accumulate a graveyard of half-broken agents nobody trusts.
“What is the most expensive mistake you have made shipping something with a model in production?”
Testing: whether they have actually run things in production, which is a completely different world from building demos.
Strong answer: a specific, slightly painful story. A cost blowout, a hallucinated output that reached a customer, a silent failure that ran for a week. The specificity is the signal.
Red flag: they have never had one. Then they have never run anything real. Everyone who has shipped models to production has a scar.
On measurement (they are marketers, not engineers)
“The last thing you built. What number did it move, and how did you know it was that thing and not something else?”
Testing: whether they think like a marketer measured on pipeline, or like an engineer measured on shipping.
Strong answer: a real metric, and an honest account of attribution. Bonus points if they volunteer the confounders instead of pretending the attribution was clean.
Red flag: the number is “the team felt faster” or “it saved time,” with no attempt to quantify. This role is measured like marketing. If they cannot connect a build to a number, they will build things that feel good and move nothing.
“How would you measure whether a research agent you built is actually good, beyond it running?”
Testing: eval thinking. Do they have a way to know their output is right, or do they ship and hope?
Strong answer: they describe a real eval. A gold set, a spot-check ritual, a human-in-the-loop sample, something. They know “it runs” and “it is good” are different claims.
Red flag: “it works” is the whole answer.
On taste, again (the artist half nobody tests)
Test this directly, with real material. Do not ask about it. Show them something and watch.
“Here is a landing page. What is wrong with it?” (Bring a real, mediocre one.)
Testing: whether they can read work and know when it is slop. This is the exact skill that separates a marketing engineer from a junior engineer with a prompt library.
Strong answer: they see it fast. The headline is doing nothing, the proof is generic, the CTA is buried. They react like someone with an editor’s eye.
Red flag: they critique the tech stack, or they cannot find anything wrong, or their fixes are cosmetic.
“Here is an AI-written outbound email. Would you send it? What would you change?”
Testing: the single most practical version of the taste test. Can they tell good output from bad output? Because their whole job is producing output at scale, and if their taste is off, the scale just multiplies the slop.
Strong answer: they can tell it is machine-written in one read, name exactly what gives it away, and rewrite the tell out of it.
Red flag: they would send it. If they cannot catch AI slop, every system they build will produce it by the thousand.
The tells you can catch in the first call
You do not need the full loop to spot most of the field. These show up in the first thirty minutes.
Red flags:
They lead with tools, not outcomes. Every answer starts with the stack. You asked what they built and they told you what they built it with. Push past the tools and see if an outcome is there.
They cannot name something they killed. No graveyard means no taste or no volume.
They need you to define the problem before they will engage. Spec-waiter. This is the disposition that does not survive contact with the role.
They bring up the newest model unprompted, twice. Hobbyist. The best ones are boring about the technology and precise about the output.
You cannot tell what number anything moved. They measured themselves like an engineer.
Green flags:
Boring about the tech, precise about the output. They do not care which model, they care whether it landed.
They volunteer what did not work. Unprompted honesty about dead tools is the strongest single signal in the loop.
They ask sharp questions about your business before pitching a single build. They are trying to find the leverage, not show you their toolbox.
They name real numbers, with the caveats attached. Real operators hedge their own attribution, because they have been burned by clean-looking numbers before. Overconfident attribution usually means they have not had to defend it.
The one kind of credential worth a second look
I told you resumes do not tell you much…and I meant it. Most marketing certifications are noise, because most of them are a quiz you can pass by watching videos at double speed. They prove someone sat through a course. They do not prove someone can do anything.
There is one exception worth knowing, and it is worth knowing precisely because it is not a quiz. Profound University runs certifications built specifically for this role, and they are structured as auditions, not tests. The Agent Engineering track makes a candidate architect and build a production-grade, multi-node agent from a real workflow in their own job (knowledge bases, structured outputs, conditional logic, real data), then present it and defend it on camera. The Analytics track makes them run a full diagnostic loop on live data, find a real opportunity, and defend the diagnosis with a root cause and a specific recommendation. Every submission is graded by hand by Profound’s team, not by an autograder.
Read that back against everything above. It is the same bar as the work sample you were about to run anyway: build something real, make it run, defend where it breaks. Which means if a candidate shows up already holding the Agent Engineering diploma, someone already made them audition for the hardest part of this job, and they passed it in front of experts.
That is the rare credential that earns a second look at the top of the funnel. Treat it as a strong green flag, not a substitute. You still run your own work sample. But it is the one line on a resume for this role that actually predicts the thing you are hiring for. If you want to point candidates at it, or size up your own bench, it lives here.
Reference checks worth the call time
References are where most loops go to confirm what they already decided. The standard questions (”what are their strengths)”, “would you hire them again”) are useless, because everyone says yes. Ask these instead.
“What did they ship that other people depended on to do their work?” If the reference cannot name one, the candidate was building for themselves.
“What did they build that nobody used, and how did they handle it?” You are checking their graveyard story against someone who watched it happen.
“Did people ask for more of what they did, or did you have to assign them work?” Pull versus push. A real marketing engineer creates demand for their own output. This is the single best reference question for this role.
“If you were starting a new team tomorrow, would this person be one of your first five hires? Why or why not?” Much harder to answer politely than “would you hire them again,” which forces a real answer.
Listen for hesitation on the pull-versus-push question specifically. A reference who has to think about whether people asked for more is telling you the work was invisible, which is the failure mode from the hire piece arriving early, for free, before you spent a dollar.
The debrief: three questions, honestly
By the end of the loop you should be able to answer three things. If you cannot, the loop was too soft and you should run the work sample again rather than voting.
Did they build something that ran, on live input, in front of you? Not described. Ran. If the answer is no, you did not test the one thing that matters and the rest of your notes are about a conversation.
Can they tell you what not to build, and why? If every answer was about what they would build and none about what they would skip, you found a builder without a filter. Call that the expensive-intern failure mode. It takes two quarters to surface and costs you both of them.
Would a marketer on your team be faster because of them, and can you point to who and at what? If you cannot picture the specific person and the specific workflow that gets better, they are not going to make anyone faster. You found a smart person who builds things.
Three yeses, make the offer and defend the comp band the way the first piece told you to. Any no, and the loop already told you where the gap is. Do not talk yourself past it. A wrong hire here gets worse the longer you leave it, in the same direction and at the same speed a right one compounds.
Steal this: the whole interview kit
That was the reasoning. Here are the scripts. Copy them into a doc, drop them in your Greenhouse, run them Monday. Swap the bracketed parts for your business. No attribution needed.
If you only run one thing, run the work sample. It catches more of the field than the rest of the loop combined.
How to sequence it
Total candidate time: about six hours, three of them paid. Total loop time: under two weeks if you schedule it up front.
1. The screener (30 minutes)
Five questions. You are listening for the tells, not grading answers yet. Two weak answers here and you stop.
Walk me through the last thing you shipped that someone else depended on to do their job. What did they use it for?
What have you built that nobody ended up using?
Tell me about something you decided not to build, and why.
When was the last time you paid for a tool instead of building it yourself? What made buying the right call?
What has your name on it and is running in production right now?
Advance if: they name outcomes before tools, they have a graveyard story ready, and at least one answer includes a number with a caveat attached.
Stop if: two or more of the red flags from the tells section show up. Most commonly it is the stack-first answer and the empty graveyard, together.
2. The work-sample brief (send this word for word)
We do not do whiteboard exercises. We want to see how you work, so we give every candidate one small problem and three hours. This is paid, and there is a hard time cap, so please do not go over. We are not looking for a finished product. We are looking for how you scope, what you cut, and whether it runs on live input.
The problem: [pick one and attach the files]
Here are ten customer call transcripts. Build something that turns them into a usable content brief.
Here is our category and the five prompts a buyer would type into an AI tool to research it. Show us how we show up in the answers today, find the gap costing us the most, and build something that starts to close it. Tell us what you would measure to know it worked.
Here is our pricing page and a competitor’s. Build a tool or agent that flags a meaningful difference and drafts a one-paragraph response.
Here is a messy CRM export. Build the enrichment and scoring pass you would run, and tell us what you would not trust in the output.
When you present, be ready to tell us: what you clarified before you started, what you cut and why, what number this would move if it were live, and where it breaks. Especially where it breaks.
Bring whatever tools you want. We do not care what you use.
Presentation structure (45 minutes): 5 min setup, 15 min they demo it running, 10 min your questions, 15 min they break their own work.
What you are grading, in order of weight:
Did it run on live input? (Pass/fail. No partial credit.)
What scoping questions did they ask before starting?
What did they cut, and could they defend the cut?
Did they narrate tradeoffs unprompted while building?
Could they name where it breaks, specifically, in the last fifteen minutes?
3. Judgment round 1: what to build, build vs buy (45 minutes)
Pick five. The first three are not optional.
Walk me through the last three things you built. For each one, tell me what you decided not to build to make room for it.
We are about to renew [real tool] at [real number] a year. Would you build the replacement? Walk me through how you decide. (Do not accept a fast answer. Push twice. The trap is that “I could build that in a weekend” sounds impressive and is wrong.)
What is something on your team that could be automated, that you would leave manual anyway? (A hypothetical counts. You are grading reasoning about ownership cost, not recall.)
Show me something you built that nobody used. What did you learn?
Tell me about a time you decided to buy instead of build. What made buying right?
Here is a snapshot of my marketing team’s week [describe a real one]. Where is the highest-return thing to build, and why not the other four?
It is your first thirty days here. What do you build first, and what do you deliberately ignore?
A stakeholder asks you to build something you think is a waste of time. What do you do?
4. Judgment round 2: reliability, measurement, taste (45 minutes)
Pick five. Number five and six are not optional and you need real material for both.
An agent you built has been running for four months. How do you know if it is still worth running?
What is the most expensive mistake you have made shipping something with a model in production?
The last thing you built: what number did it move, and how did you know it was that and not something else?
How would you measure whether a research agent you built is any good, beyond the fact that it runs?
[Show a mediocre landing page.] What is wrong with this?
[Show an AI-written outbound email.] Would you send this? What would you change?
Show me the ugliest thing you ever shipped that worked. Why was ugly the right call?
5. Reference questions (two calls)
What did they ship that other people depended on to do their work?
What did they build that nobody used, and how did they handle it?
Did people ask for more of what they did, or did you have to assign them work? (The one that matters. Listen for hesitation, not the answer.)
If you were starting a new team tomorrow, would this person be one of your first five hires? Why or why not?
6. The scorecard (paste into Greenhouse)
Score 1 to 4 on each. A single 1 on Judgment, Build, or Taste is a no regardless of the rest.
Judgment, knows what to build and what to skip 1: builds whatever is asked · 2: prioritizes when told the goal · 3: finds the return themselves · 4: tells you what not to build and is right
Build, it runs 1: demos only, nothing survives live input · 2: ships, but fragile · 3: ships things that run in production · 4: ships things others build on top of
Taste, tells good output from slop 1: cannot spot slop · 2: spots it when pointed at it · 3: catches it fast, fixes the tell · 4: their default output is already clean
Ownership, thinks past the ship date 1: ships and forgets · 2: monitors when reminded · 3: watches cost, drift, and failure by habit · 4: kills their own dead tools before you notice
Marketer instinct, connects builds to numbers 1: measures in “felt faster” · 2: names a metric, weak attribution · 3: ties builds to numbers with honest caveats · 4: chooses what to build by the number it will move
Decision rule: three or more 4s, and no 1s on Judgment, Build, or Taste. Anything else, walk.
Calibration notes for the debrief. Two things go wrong in the room and both are predictable.
The first is that someone will argue a 2 on Build up to a 3 because the candidate was impressive to talk to. Build is the one score that is not a matter of opinion. It ran on live input or it did not, and if nobody can point to the moment it ran, it is a 2.
The second is that a candidate with 4s everywhere and a 1 on Taste will get an offer anyway, because four out of five feels like a strong loop. It should not. Taste is the score that predicts the failure mode you cannot see for two quarters, which is a person who ships constantly and whose output nobody wants. Walk on the 1. You will not regret it and you will never find out you were right, which is what makes it hard.
The line
Anyone can learn the tools this year. Which is exactly why the tools are the wrong thing to test for.
The candidate who impresses you by naming every platform is easy to be dazzled by and usually the wrong hire. The one worth hiring is quieter about the stack, faster to ask what your business needs, honest about the things they built that died, and unable to stop themselves from telling you what they would not build. You will not find that person with better questions about their resume. You will find them by making them work in front of you and watching what they cut.
Your first marketing engineer sets the bar for the next five. Get the vetting right and the bar is judgment. Get it wrong and the bar is whoever interviewed best, which in this role is the wrong person almost every time.
You have the req. Now you have the loop.
Go find a real one.




