Add CLAUDE.md, project memory notes, and token_count
CLAUDE.md points future sessions at .claude/memory/ for accumulated project context (purpose, SA/MR topology, packaging/signing decisions, open external dependencies) not fully captured elsewhere in the repo.
This commit is contained in:
@@ -0,0 +1,6 @@
|
|||||||
|
# Memory Index
|
||||||
|
|
||||||
|
- [Project Purpose](project-purpose.md) — why this repo exists, the two core problems being solved
|
||||||
|
- [SA/MR Topology](topology.md) — deployment topology terminology and shape
|
||||||
|
- [Packaging & Signing Design Decisions](packaging-signing-decisions.md) — key architecture calls and why
|
||||||
|
- [Open External Dependencies](open-external-dependencies.md) — things this design assumes but doesn't own
|
||||||
@@ -0,0 +1,15 @@
|
|||||||
|
# Open External Dependencies
|
||||||
|
|
||||||
|
Things the packaging/signing design assumes exist but doesn't own or
|
||||||
|
fully specify — needed before implementation:
|
||||||
|
|
||||||
|
- **FEC HTTP service** for the MR one-way hop — assumed already available,
|
||||||
|
~1GB chunks, not built as part of this design.
|
||||||
|
- **Existing S3 handover interface spec** — a separate, already-established
|
||||||
|
convention elsewhere in this environment for mapping S3 user object
|
||||||
|
metadata (`x-amz-meta-*` headers) to NiFi attributes via a regular
|
||||||
|
expression, plus a flag indicating whether the attribute set exceeded
|
||||||
|
S3's ~2KB metadata limit (in which case it falls back to this design's
|
||||||
|
JSON trailer). The exact regex pattern and flag/tag name are **not**
|
||||||
|
reproduced in the packaging/signing spec — need to source that document
|
||||||
|
before implementation.
|
||||||
@@ -0,0 +1,39 @@
|
|||||||
|
# Packaging & Signing — Key Decisions
|
||||||
|
|
||||||
|
Full detail lives in
|
||||||
|
`docs/superpowers/specs/2026-07-28-flowfile-packaging-signing-design.md`.
|
||||||
|
These are the non-obvious calls worth remembering, several of which
|
||||||
|
reversed an earlier direction during design review:
|
||||||
|
|
||||||
|
- **No sidecar fronting ingestion.** A claim-check sidecar to bypass
|
||||||
|
NiFi's content repo for large payloads was considered and rejected —
|
||||||
|
all content must transit NiFi for routing/content-assurance controls.
|
||||||
|
The external service's role is post-NiFi only: assembly, stall
|
||||||
|
monitoring, cleanup, catalog DB maintenance.
|
||||||
|
- **No external signing service.** Vault Transit / a network-isolated
|
||||||
|
signer was considered and rejected in favor of in-process signing
|
||||||
|
(Bouncy Castle OpenPGP, key loaded via a Controller Service) to keep the
|
||||||
|
architecture self-contained within NiFi. This is a deliberate
|
||||||
|
non-repudiation trade-off, stated explicitly in the spec — key custody
|
||||||
|
now depends on host-level protection (sensitive-property encryption,
|
||||||
|
restricting which instances have the Controller Service) rather than
|
||||||
|
network isolation.
|
||||||
|
- **Clear-sign (OpenPGP), not a bespoke signature format.** Chosen
|
||||||
|
specifically because an operator on the far side of MR can verify an
|
||||||
|
artifact by hand with stock `gpg --verify` and a public key — no custom
|
||||||
|
SDK needed. Matters because MR already has a manual-operator recovery
|
||||||
|
model for missing chunks (adopted from the Phat Files design).
|
||||||
|
- **HMAC was considered and rejected** as the integrity mechanism — it
|
||||||
|
detects tampering but can't give non-repudiation, since every verifier
|
||||||
|
necessarily holds the same shared secret and could in principle have
|
||||||
|
forged the tag themselves.
|
||||||
|
- **Content + JSON trailer as one object, not two.** Chosen for
|
||||||
|
end-to-end loss protection: two separate objects (content, manifest)
|
||||||
|
have independent failure/lifecycle/replication paths that can diverge.
|
||||||
|
A single completed write can't end up in a split state.
|
||||||
|
- **S3 attribute handover defers to an existing, established handover
|
||||||
|
interface spec**, not this design's own invention — see
|
||||||
|
[[open-external-dependencies]].
|
||||||
|
- **S3-vs-XFS storage backend split by file size is unresolved** — open
|
||||||
|
question whether S3 can serve small files efficiently enough to drop
|
||||||
|
the XFS tier entirely.
|
||||||
@@ -0,0 +1,22 @@
|
|||||||
|
# Project Purpose
|
||||||
|
|
||||||
|
This repo is about improving NiFi workflows in two related ways:
|
||||||
|
|
||||||
|
1. **Performance** — the current pack/unpack path (`MergeContent` /
|
||||||
|
`UnpackContent`) is disk-I/O bound on shared vSAN storage, due to
|
||||||
|
redundant read/write passes: an ingest write, a separate read to hash
|
||||||
|
content, a read+write to merge/pack, a read+write to unpack. vSAN is
|
||||||
|
shared, virtualized IOPS across the vSphere cluster, so the pain is
|
||||||
|
*contention*, not any single slow disk.
|
||||||
|
2. **Integrity & non-repudiation** — FlowFiles packed and exported outside
|
||||||
|
NiFi's space need signed attributes (the attributes already carry a
|
||||||
|
content hash, so content isn't re-hashed at pack time) to prove both
|
||||||
|
integrity and authorship, for S3 storage/retrieval, HTTP/filesystem
|
||||||
|
delivery, and selective date/category replay from a catalog database.
|
||||||
|
|
||||||
|
A future requirement extends this to TB-scale files that must be chunked,
|
||||||
|
tracked, and transited through NiFi (for routing/content-assurance
|
||||||
|
controls — all content must pass through NiFi) — including across a
|
||||||
|
one-way network boundary (MR) with no return channel for retransmission.
|
||||||
|
|
||||||
|
See [[topology]] and [[packaging-signing-decisions]].
|
||||||
@@ -0,0 +1,27 @@
|
|||||||
|
# SA/MR Deployment Topology
|
||||||
|
|
||||||
|
**SA** is the name of the whole system pictured in `grand_plan.jpg`: two
|
||||||
|
chains of independent, single-function NiFi hosts (each host linked to the
|
||||||
|
next via Site-to-Site) bridged in the middle by **MR**.
|
||||||
|
|
||||||
|
**MR** is the one-way relay (diode-style boundary) between the two chains
|
||||||
|
— forward error correction via an HTTP service, ~1GB target chunk size, no
|
||||||
|
return channel for retransmission requests.
|
||||||
|
|
||||||
|
Content flows one direction through SA: external S3/S2S in -> first NiFi
|
||||||
|
chain (node 1 -> 2 -> 3) -> MR -> second NiFi chain (node 1 -> 2 -> 3) ->
|
||||||
|
external S3/S2S out. The Replay DB branches off **before** MR, via a
|
||||||
|
"copy" from node 1 of the *first* chain — i.e. at the true ingest point
|
||||||
|
into SA, not after delivery. This was a real correction during design
|
||||||
|
review — an earlier draft had the Replay DB positioned after MR, which was
|
||||||
|
wrong.
|
||||||
|
|
||||||
|
Chunking and hashing both happen at that same ingest point (node 1 of the
|
||||||
|
first chain), in a single streaming pass, so the same chunked
|
||||||
|
representation transits both the intra-chain S2S hops and the MR relay
|
||||||
|
uniformly — no separate chunking scheme per transport, and no size
|
||||||
|
threshold decides whether to chunk (a small file is just the one-chunk
|
||||||
|
case).
|
||||||
|
|
||||||
|
Source: whiteboard sketch `grand_plan.jpg`, reconciled into
|
||||||
|
`docs/superpowers/specs/2026-07-28-flowfile-packaging-signing-design.md`.
|
||||||
@@ -0,0 +1,45 @@
|
|||||||
|
# CLAUDE.md
|
||||||
|
|
||||||
|
Project notes for the NiFi prototyping repo — improving NiFi workflow
|
||||||
|
performance, packaging/signing of FlowFiles, and large-file (TB-scale)
|
||||||
|
transfer across a one-way network boundary.
|
||||||
|
|
||||||
|
## Before starting work
|
||||||
|
|
||||||
|
Read everything in `.claude/memory/` first — it holds accumulated notes,
|
||||||
|
terminology, and context (some of it non-obvious from the code/docs alone)
|
||||||
|
that this file only summarizes. Treat it as living project memory: update
|
||||||
|
it as significant new decisions or context emerge, don't just accumulate.
|
||||||
|
|
||||||
|
## Key documents
|
||||||
|
|
||||||
|
- `2026-07-23-phat-files-design.md` — approved spec for chunked TB-scale
|
||||||
|
S3-to-S3 transfer across a one-way boundary. Source of proven concepts
|
||||||
|
(chunking, resend/duplicate handling, server-side assembly) reused by the
|
||||||
|
packaging/signing design below — not a dependency of it.
|
||||||
|
- `docs/superpowers/specs/2026-07-28-flowfile-packaging-signing-design.md`
|
||||||
|
— design for FlowFile packaging (content + JSON trailer), in-process
|
||||||
|
OpenPGP signing, uniform chunking, and S3/XFS storage backend handling.
|
||||||
|
Status: draft, pending user review.
|
||||||
|
- `grand_plan.jpg` — whiteboard sketch of the SA/MR deployment topology the
|
||||||
|
design above was reconciled against.
|
||||||
|
|
||||||
|
## Terminology
|
||||||
|
|
||||||
|
- **SA** — the overall system: two chains of independent, single-function
|
||||||
|
NiFi hosts (linked via Site-to-Site within each chain) bridged by MR.
|
||||||
|
- **MR** — the one-way relay/diode boundary between the two NiFi host
|
||||||
|
chains. No return channel; forward error correction via an HTTP service,
|
||||||
|
~1GB target chunk size.
|
||||||
|
|
||||||
|
## Environment
|
||||||
|
|
||||||
|
- On-prem VMs, vSAN-backed storage — shared IOPS across the vSphere
|
||||||
|
cluster, so contention (not per-VM disk speed) is the bottleneck.
|
||||||
|
- NiFi runs as independent single-node instances (no clustering), each
|
||||||
|
with a distinct logical function, scaled horizontally.
|
||||||
|
|
||||||
|
## Remote
|
||||||
|
|
||||||
|
- `https://gitea-5001.taild45a1b.ts.net/Hackathon/GrandPlan` (branch
|
||||||
|
`main`) — credentials supplied out-of-band, not stored in this repo.
|
||||||
@@ -0,0 +1 @@
|
|||||||
|
current_count=250K, model=sonnet5, effort=high
|
||||||
Reference in New Issue
Block a user