UCB — Hard Technical Questions

Diagrams carry the structure; text is the spoken answer · Sources: cheat-sheet library + Moses ADRs, refreshed Oct 2026 · Landscape, prints to two pages

Agents"How do you build agents?"

Orchestratorowns state · routes tasks
↓
Agent Astateless · idempotent
Agent Bstateless · idempotent
Agent Cstateless · idempotent
tools via MCP · every call returns result + confidence + failure reason
↓
Reflection gatereviews output before delivery

From primitives, not frameworks — 2025's framework wars resolved toward thin orchestration over model APIs. Failure taxonomy is known (emergent composition failures, hallucination compounding through chains, wrong tool selection, reasoning loops) — so agents are tested like software: decompose the trace, unit-test sub-tasks against ground truth, full-chain regression after any prompt change.

Agents"What types of agents do you use?"

Pipelineingest → summarize → score · 7,000 briefs/day
JudgeClaude Sonnet 4.5 LLM-as-judge + HHEM/BERTScore — a generator can't verify its own output
Persona simulationblinded 4-clinician QA harness · inter-rater agreement
CodingClaude Code · Kimi Code · OpenCode on local vLLM — delivery leverage

Commercial mapping: Veeva's pre-call and voice agents are pipeline + reflection, productized. The orchestration layer above them is where SKAI sits.

KG"How do you build out a knowledge graph?"

Build lane — deterministic first
Crossref
OpenAlex
→
zero-LLM bulk insertcitation · authorship · affiliation edges
→
Postgres — canonicalkg_node / kg_edge · closed vocab · provenance · pgvector
one transaction: node + vector + source
Query lane — semantic edges deferred
Query
→
adaptive routerinstrumented escalation
→
≥70% vector tierrecursive CTE, 2–3 hop cap
else semantic edgesLazyGraphRAG-style

Graph is a disposable projection, not a second store — Kuzu archived 2025, Apache AGE analyst-sugar only. Qualifiers (dose, population, timeframe) ride as hyper-relational facts, not flattened triples.

RAG"How do you do RAG / vectors in Postgres?"

Query
→
BM25Postgres full-text
Densepgvector · HNSW over halfvec
limits: 2,000d vector · 4,000d halfvec
→
RRF fusereciprocal rank fusion
→
cross-encoder rerankQwen3 0.6B
→
LLM + verifytwo-tier scoring

Hybrid from day one — dense-only retrieval lost that fight. Embeddings: Qwen3-Embedding 4B/8B on MLX when data control matters, BGE-M3 held as baseline — every row persists model revision, dimensions, normalization, chunk policy, source hashes; parallel namespaces per encoder, never mixed. PG19 adds SQL/PGQ property-graph views — read-only sugar. Measured: recall@k, evidence precision, citation validity, unsupported-claim rate.

Decision"RAG or fine-tune?"

RAG when…knowledge changes often · traceability required · hallucination liability · governed content (MLR-approved chunks)
Fine-tune when…stable format & vocabulary · latency/cost — distill teacher → small model

At UCB the answer starts RAG over governed Vault content; fine-tuning is a later cost play, not a starting point.

Eval"How do you evaluate agents at scale?"

Machine scoring — 500+ records/hrHHEM faithfulness + BERTScore · two-tier, because one metric lies
↓
Blinded persona harness4 clinician personas · inter-rater agreement
↓
SME parallel run — gate88% blinded preference before production

Earned caution: HHEM alone ran near-chance on some failure modes and failed silently — scoring is always two-tier. Agent-written code gets human-grade governance: doc-lint CI, DDL-parity gates, required review.

Models"Why Qwen over GPT or Claude?"

Bake-off
rankings · benchmarks · own evals
→
Fit to job
faithfulness · context · cost · latency
→
Deploy
vLLM · MLX · API — swappable

Selection is a methodology, not a brand: Artificial Analysis rankings, SWE-rebench, LMArena, then own bake-offs against the gold corpus. Model-agnostic interfaces so the winner is swappable — the same discipline decides what runs on Azure OpenAI at UCB.

Oct 2026 refresh vs. the Aug 2026 sheets: models moved a generation (Qwen3.6, Claude Sonnet 4.5) · MCP is now the tool standard · Veeva AI agents shipped Dec 2025 on Bedrock — UCB's Azure tension · agent frameworks consolidated — never lead with LangChain · GraphRAG / hyper-relational moved from research to shipped. Say-this-not-that: never say "POC" without "which became production at N." · never say "I'd use LangChain for that."