Diagrams carry the structure; text is the spoken answer · Sources: cheat-sheet library + Moses ADRs, refreshed Oct 2026 · Landscape, prints to two pages
From primitives, not frameworks — 2025's framework wars resolved toward thin orchestration over model APIs. Failure taxonomy is known (emergent composition failures, hallucination compounding through chains, wrong tool selection, reasoning loops) — so agents are tested like software: decompose the trace, unit-test sub-tasks against ground truth, full-chain regression after any prompt change.
Commercial mapping: Veeva's pre-call and voice agents are pipeline + reflection, productized. The orchestration layer above them is where SKAI sits.
Graph is a disposable projection, not a second store — Kuzu archived 2025, Apache AGE analyst-sugar only. Qualifiers (dose, population, timeframe) ride as hyper-relational facts, not flattened triples.
Hybrid from day one — dense-only retrieval lost that fight. Embeddings: Qwen3-Embedding 4B/8B on MLX when data control matters, BGE-M3 held as baseline — every row persists model revision, dimensions, normalization, chunk policy, source hashes; parallel namespaces per encoder, never mixed. PG19 adds SQL/PGQ property-graph views — read-only sugar. Measured: recall@k, evidence precision, citation validity, unsupported-claim rate.
At UCB the answer starts RAG over governed Vault content; fine-tuning is a later cost play, not a starting point.
Earned caution: HHEM alone ran near-chance on some failure modes and failed silently — scoring is always two-tier. Agent-written code gets human-grade governance: doc-lint CI, DDL-parity gates, required review.
Selection is a methodology, not a brand: Artificial Analysis rankings, SWE-rebench, LMArena, then own bake-offs against the gold corpus. Model-agnostic interfaces so the winner is swappable — the same discipline decides what runs on Azure OpenAI at UCB.