For hiring teams

Your FDE opening is about to get buried in applicants. We grade which ones can actually deploy.

The New York Times covered the Forward Deployed Engineer role on 2026-08-30, citing Indeed data showing listings up 730 percent in a year. Every posting now draws a flood, and a resume cannot show you the one thing the job is: judgment under a customer's real constraints.

LeetCode does not test it either. A deployment case does. We run timed, realistic customer scenarios and grade every answer against the rubric the role actually needs, so your team interviews the top five instead of phone-screening fifty.

How it works

  1. You send each applicant a private case link: a timed 45 to 90 minute deployment scenario, picked to match your role.
  2. Their written answer is graded against a six-dimension FDE rubric, with every hidden trap they caught or missed called out.
  3. You get a ranked sheet plus a per-candidate graded report you can bring straight into the debrief.

The cases

Five scenarios, each built like a real first customer call, each with hidden traps that separate practitioners from people who have read about agents:

  • Enterprise support agent, healthcare. An L1 agent for a 14-hospital system that must never invent clinical advice. Tests scope surgery under safety constraints.
  • AP invoice automation, finance ops. An agent in the payables loop where a wrong payment is unrecoverable. Tests controls, evals and human-in-the-loop design.
  • Company-wide AI assistant, enterprise IT. A work assistant across permissioned internal systems. Tests access boundaries and grounding discipline.
  • Coding-agent rollout, engineering org. Rolling agents into a 400-engineer org. Tests eval design, canary strategy and the politics of developer tooling.
  • Multi-agency data intelligence, public sector. Case intelligence across agencies that do not share schemas or trust. Tests discovery and constraint mapping at its hardest.

The rubric

Every answer is scored 1 to 5 on each dimension, with a one-line justification per score:

  • Discovery and constraint mapping
  • Intent and scope surgery
  • Agent architecture
  • Eval plan
  • Rollout and stakeholder design
  • The asks: what they demand from you before building

What a report looks like

Sample report for the healthcare case. This candidate reads well on paper; the rubric shows exactly where they would hurt you in production:

SCORE: 3/5

Discovery and constraint mapping: 4/5. Inventoried PHI exposure, Epic read-only limits and the clinical-refusal requirement before proposing anything. Missed multi-hospital policy variance: the answer assumes one policy source where fourteen hospitals have their own.

Intent and scope surgery: 2/5. Correctly cut clinical advice from v1 and picked password resets plus reschedules. But the volume math sums raw intent shares to claim the 40 percent deflection target is reachable. With read-only integrations, write-side reschedules still escalate; the honest containable number is roughly half the claim. This is the difference between a pilot that survives its board review and one that does not.

Agent architecture: 4/5. Tool-calling over Epic and Zendesk, policy retrieval with citations, hard escalation rules. Kept voice out of the pilot, correctly.

Eval plan: 3/5. Proposed golden sets and named hallucinated clinical advice as a P0 failure class. Did not notice that no intent taxonomy exists yet, so the golden sets have nothing to be built from; the plan starts one step later than reality allows.

Rollout and stakeholder design: 3/5. Shadow mode before canary, kill switch, sensible metrics. Treated the marketing team's already-announced launch date as a deadline to hit rather than a conflict to surface, and never asked about the BPO contract terms that decide whether the savings are real.

The asks: 3/5. Asked for ticket exports and an Epic sandbox. Did not ask who owns policy content, which is the question that decides whether grounding is buildable at all.

Biggest gap: numeric honesty under constraints. The candidate presents intent-share sums as a deflection forecast. In front of your customer, that becomes a commitment your deployment cannot keep.

Pricing

Pilot: your first five candidates, free. After that, $49 per graded candidate. Volume and custom cases by conversation. No platform to adopt, no seats, no contract; you send candidates, you get reports.

Who is behind this

Armando Gonzalez

Armando Gonzalez. Two years building AI systems in production, forward-deployed work included. These cases and this rubric run A10X's training arena, where engineers preparing for the FDE role drill the same scenarios. Screening is the same engine pointed at your pipeline. I also run embedded FDE engagements when you need the work done, not just the hire.

Start a pilot: hello@a10x.dev

Include a link to your role and roughly how many candidates you have. First graded reports typically within 48 hours of your candidates submitting.