SHEPHERD ELECTRIC · CMMC + AI UPDATE · 2026-08-28

The internal layer
is energized

The CMMC route holds. The owned-AI layer is running code, not strategy. And the hardware to scale it now has ship dates.

Will · prepared for Billy · trim before Nardo / Earl

TL;DR · three things converged this month

Same plan. New column: delivered

HOLDING

The CMMC route

Unchanged since 08-03. Entra identity, one Eclipse gateway, governed Bedrock, certified S3. The internal layer slots inside this spine — it replaces nothing.

LIVE

The owned layer

Open-weights Qwen 3.8 serving on our own silicon today, fronted by an MIT-licensed agent harness we audited line-by-line and already re-branded once.

DATED

The hardware

Apple's M5 Ultra Mac Studio: 256GB ships Sept 22, 512GB late October. The "<1 year to a local stack" promise is now beatable by months.

THE SPINE · four feeders, one bus · per 08-03 plan

The CMMC one-line

Entra

IA / AC

Identity. SSO, MFA, offboarding — Graybar tenant. The 07-30 business case is the artifact.

Eclipse GW

AC / SC

ERP boundary. One authenticated door; the public write path gets killed.

Bedrock

Governed AI

Cloud lane. Region-locked us., logged, no training on data, FedRAMP CUI-fallback.

S3

MP / SC

Storage. Company account, replicated, 21,714-object SHA-256 cert — delivered 07-27.

COMMON BUS — NAMED IDENTITY · MFA · ONE AUDIT LOG

The internal layer lands as the on-prem lane of the AI feeder — same bus, same log. Not a new initiative.

INTERNAL LAYER 1 of 2 · the model

Qwen 3.8 on our silicon

qwen3.8-27b (q8) serves right now on the Mac Studio — OpenAI-compatible API, reasoning model, vision-capable.

The API key lives on the box and never leaves it. That's the CMMC-relevant property: a physical boundary, not a vendor's policy promise.

Rule of the layer: sensitivity picks the rails; capability picks the model within rails.

LIVE ●
serving today · llama-server :8080
$0
model license — open weights (Qwen / Apache-2.0 line)
0 bytes
of payload leave the building at inference time
INTERNAL LAYER 2 of 2 · the harness

DeepSeek Harness: MIT, audited, ownable

CheckWhat we verified — by reading the code, not the README
LICENSEMIT. Fork it, theme it, ship it. No vendor to ask.
LOCAL MODELConfig, not code. Our llama-server plugs in as an OpenAI-compatible provider route — Qwen on the Studio is a supported backend today.
NO LOCK-INThe agent loop itself is replaceable without forking — a documented factory seam, one config row. Everything is a plugin.
AUDIT TRAILAppend-only session log, built-in compaction with provenance — the record shape a CMMC reviewer wants from an AI system.
PROOFAlready re-branded once (08-25): a fully themed copy shipped for another client. A "Shepherd Harness" is the same copy-and-theme operation.

KNOWN LIMITNo per-request constrained decoding — structured output rides a validated tool call. Acceptable; on record so nobody re-discovers it.

WHY IT MATTERS · the questions Earl's program asks every vendor

Every vendor gets asked. We can answer.

CMMC questionInternal-layer answer
DATAWhere does it go? Nowhere. Inference on owned hardware; key never leaves the box.
ACCESSBehind the same Entra identity spine as everything else. One bus.
AUDITAppend-only session log + per-turn provenance stamps (local vs cloud, every turn).
GOVERNANCEWeights on our disk, versions we pin, a harness we can read line-by-line.
LOCK-INNone — model-agnostic by architecture. Verified, not claimed.

Any purchased AI product faces these same questions — with a vendor's paper instead of our code.

HARDWARE · Apple M5 Ultra Mac Studio · announced 08-25

The on-prem tier gets ship dates

ConfigShipsPrice signalWhat it runs (approx.)
M5U · 96GBSept 22from $5,499Current Qwen line with huge headroom; ~70B-class q4
M5U · 256GBSept 22≈ $9,499 First Shepherd-grade inference box. 200B+-class MoE, quantized.
M5U · 512GBLate Octno price publishedDeepSeek-class 671B MoE ~q4 (~380–400GB) — fully on-prem frontier
1.2 TB/s
memory bandwidth — the binding constraint for tokens/sec
No DC
ordinary capital purchase · plugs into the wall · inside the enclave boundary

CAUTIONApple has not priced the 512GB config. Nothing in writing with a 512 number until they do.

THE NARRATIVE · what moves, what holds

Changed / unchanged

● Changed

  • "Own the weights" moved to the delivered column — running endpoint, audited harness, themed proof.
  • The <1-year local-stack promise is now conservative — a live demo is assemblable the week the 256GB box arrives.
  • Buy-vs-build gains a floor: $0 harness + $0 weights + ≈$9.5K box — less than most vendor pilots' first quarter.

○ Unchanged

  • Bedrock stays the governed cloud path for CUI-shaped work. The local lane is additive, not a replacement.
  • Right-size rule holds. Two people + Earl + Nardo. Small asks only.
  • The cutover still ships first. Zero scope added to DTS-on-AWS.
NEXT ACTIONS · four items, all small

What happens next

1 · DEMO30-min Billy walkthrough: Qwen through the harness, the re-brand proof, decide when a "Shepherd Harness" theme pass happens.
2 · DECIDEHardware call by ~Sept 15: 256GB now (low-regret if the demo lands) vs wait for 512GB pricing. Decide on evidence.
3 · FOLD INOne paragraph into the Earl/Nardo package: "the on-prem lane of the governed-AI feeder." Same spine, same bus.
4 · SHOWThe weights workshop now has a live exhibit — a model you can point at, on a box in the room.

Let's GrOw.

F fullscreen · ← → navigate
CKT 01 / 09