The CMMC route holds. The owned-AI layer is running code, not strategy. And the hardware to scale it now has ship dates.
Will · prepared for Billy · trim before Nardo / Earl
Unchanged since 08-03. Entra identity, one Eclipse gateway, governed Bedrock, certified S3. The internal layer slots inside this spine — it replaces nothing.
Open-weights Qwen 3.8 serving on our own silicon today, fronted by an MIT-licensed agent harness we audited line-by-line and already re-branded once.
Apple's M5 Ultra Mac Studio: 256GB ships Sept 22, 512GB late October. The "<1 year to a local stack" promise is now beatable by months.
Identity. SSO, MFA, offboarding — Graybar tenant. The 07-30 business case is the artifact.
ERP boundary. One authenticated door; the public write path gets killed.
Cloud lane. Region-locked us., logged, no training on data, FedRAMP CUI-fallback.
Storage. Company account, replicated, 21,714-object SHA-256 cert — delivered 07-27.
The internal layer lands as the on-prem lane of the AI feeder — same bus, same log. Not a new initiative.
qwen3.8-27b (q8) serves right now on the Mac Studio — OpenAI-compatible API, reasoning model, vision-capable.
The API key lives on the box and never leaves it. That's the CMMC-relevant property: a physical boundary, not a vendor's policy promise.
Rule of the layer: sensitivity picks the rails; capability picks the model within rails.
| Check | What we verified — by reading the code, not the README |
|---|---|
| LICENSE | MIT. Fork it, theme it, ship it. No vendor to ask. |
| LOCAL MODEL | Config, not code. Our llama-server plugs in as an OpenAI-compatible provider route — Qwen on the Studio is a supported backend today. |
| NO LOCK-IN | The agent loop itself is replaceable without forking — a documented factory seam, one config row. Everything is a plugin. |
| AUDIT TRAIL | Append-only session log, built-in compaction with provenance — the record shape a CMMC reviewer wants from an AI system. |
| PROOF | Already re-branded once (08-25): a fully themed copy shipped for another client. A "Shepherd Harness" is the same copy-and-theme operation. |
KNOWN LIMITNo per-request constrained decoding — structured output rides a validated tool call. Acceptable; on record so nobody re-discovers it.
| CMMC question | Internal-layer answer |
|---|---|
| DATA | Where does it go? Nowhere. Inference on owned hardware; key never leaves the box. |
| ACCESS | Behind the same Entra identity spine as everything else. One bus. |
| AUDIT | Append-only session log + per-turn provenance stamps (local vs cloud, every turn). |
| GOVERNANCE | Weights on our disk, versions we pin, a harness we can read line-by-line. |
| LOCK-IN | None — model-agnostic by architecture. Verified, not claimed. |
Any purchased AI product faces these same questions — with a vendor's paper instead of our code.
| Config | Ships | Price signal | What it runs (approx.) |
|---|---|---|---|
| M5U · 96GB | Sept 22 | from $5,499 | Current Qwen line with huge headroom; ~70B-class q4 |
| M5U · 256GB | Sept 22 | ≈ $9,499 | First Shepherd-grade inference box. 200B+-class MoE, quantized. |
| M5U · 512GB | Late Oct | no price published | DeepSeek-class 671B MoE ~q4 (~380–400GB) — fully on-prem frontier |
CAUTIONApple has not priced the 512GB config. Nothing in writing with a 512 number until they do.
| 1 · DEMO | 30-min Billy walkthrough: Qwen through the harness, the re-brand proof, decide when a "Shepherd Harness" theme pass happens. |
| 2 · DECIDE | Hardware call by ~Sept 15: 256GB now (low-regret if the demo lands) vs wait for 512GB pricing. Decide on evidence. |
| 3 · FOLD IN | One paragraph into the Earl/Nardo package: "the on-prem lane of the governed-AI feeder." Same spine, same bus. |
| 4 · SHOW | The weights workshop now has a live exhibit — a model you can point at, on a box in the room. |
Let's GrOw.