
Applied-AI engineer who owns Provenance's intelligence layer end to end — the attribution engine and the Assistant — while the founder covers the record layer, the frontend, and day-to-day product. This is deep work in a real system, not greenfield, and it's backend/infra only (no frontend). The person needs to be in seat by mid-October; the six-month plan sequences around them.
The mandate is sequenced, not parallel. Week one: two switch-sized Assistant fixes ship behind existing telemetry — turn the grounding verifier back on and raise the throttled tool-round budget. Month one: finish the deck-attribution leg to demo-grade (it already agrees 120/120 where PowerPoint declares the link). From there: the long pole — rebuild document-attribution recall (stuck at ~0.30 against a 1.0 bar because there's no semantic retrieval, and precision-protecting binary gates delete true positives), converting binary gates into calibrated scores to recover true matches without giving back precision, while growing the Assistant's eval corpus (8 cases today → 150+). The core mission: take recall from 0.30 → ≥0.6 at precision ≥0.9 on expanded labeled sets, make "every number names its cell" real, get the Assistant to analyst-grade (reliability ≥0.9), mounted inside Excel, with a first human-approved write path. Beyond six months, this grows into ownership of the whole intelligence layer.
What You'll Own
- The attribution engine — document→cell and cell→slide lineage, retrieval, matching, entity resolution, and the calibrated scoring that replaces today's binary gates
- The Assistant — the agent loop, eval corpus, grounding verification, latency, Excel mounting, and a first cited-evidence write path
- The quality bars themselves — labeled gold sets, CI-asserted eval gates, and a live production-quality dashboard the company can sell against
- Operations — pushing improvements automatically across ~400 existing repositories (recomputation is manual today)
- Ramp (from the plan): wk 1–2 grounding on + budget lifted + funnel instrumentation; mo 1 deck leg to demo-grade, labeled cells 126→300; mo 2 semantic retrieval measured, gates redesigned as scores, corpus →50 with failure trajectories; mo 3 recall ≥0.5 @ precision ≥0.9, quality dashboard live, Assistant reliability ≥0.85 at 100+ cases
Requirements
- Has shipped both halves: retrieval/matching quality (search relevance, ranking, record linkage, document AI) and an LLM product with a real eval harness and telemetry — this combined profile is the whole point, and rarer than either half alone
- Has personally built a precision/recall or reliability harness before the feature and can tell that story with numbers (the single strongest signal)
- Treats candidate generation and selection as separate problems with separate metrics; thinks about denominators ("what fraction can even have a source?")
- Treats LLMs as calibrated components inside a measured system, not as the system
- Python; provider-agnostic instincts (Gemini-on-Vertex today — provider preference irrelevant)
- Comfortable owning a large legacy hot path (a 6,033-line matcher) under an eval gate without demanding a rewrite
Execution and Ownership
- Genuine ownership — "feels it," not a salaried ticket-taker; grows into VP of Technology
- High transparency and speed — fast async communication (Arda's top two words: "transparency and speed")
- Has built something adjacent/similar before — no learning-this-from-scratch, given the timeline
- Chip on the shoulder; wants to prove themselves; builder energy over prestige-seeking
- Adaptable, chill, low-drama; won't jump ship when things get choppy
Background
- Experience level flexible — a hungry builder with the specific spike is welcomed over an expensive generalist; the brief screens hardest for the retrieval-evaluation discipline (the half that can't be taught quickly)
- Ex-founder, or unhappy-at-current-job builder who'd rather build at a startup; strong young talent (e.g., AI-club leaders) in play if the spike is real
- Finance exposure is a plus, not a requirement
Anti-Signals (screen out)
- Generic senior full-stack engineers — this is a specialist seat
- "Just add RAG / a bigger model" as the first answer to recall; prompt-tinkerers whose quality story is "we iterated until it felt right"
- Framework maximalists who want an orchestration layer on day one (the architecture is deliberately one simple loop with deterministic guardrails)
- Research-only or demo-driven profiles with no production telemetry story; anyone uncomfortable hearing "recall 0.30" said out loud
Who Will Thrive Here
- An applied-AI engineer who's lived through vibes-driven AI roadmaps and craves labeled gold sets and CI-asserted quality bars
- Someone who wants total ownership of a moat — algorithm, evals, operations, and the number on the dashboard
- Precision engineering where precision matters — users move billions on these documents; a wrong citation is a fireable event for them
- A builder joining at ~zero revenue with a working record-layer product and a six-month line to $1M ARR
No benefits listed yet.