|
ProtoCore v1.0.16
Deterministic, zero-heap network stack for embedded targets
|
How bytes and control flow move between the OSI-style layers and who owns each cross-layer concern. This is the map referenced by the internal-piping cleanup: the rule is one owner per cross-layer concern, behind a clean API - no layer reaches into another's internals, and a function that needs a resource asks that resource's owner for it. Both axes are now settled: the data-piping axis (who owns RX / TX / window / events / memory / streaming / client I/O) and the dispatch axis (how each protocol attaches) both meet the one-uniform-seam rule, with no protocol special-cased. The piping is straight; what remains below is the map, not a to-do.
The server is a 2-thread system and every cross-layer hazard lives on one of two boundaries:
lowlevel_recv_cb copies an inbound segment into the connection's RX ring and posts a TcpEvt; lowlevel_sent_cb nudges the owning worker; listener_accept_cb assigns a new slot its owner worker.service_once() -> server_tick() (drain the event queue -> dispatch_event -> ProtoHandler::on_data) then a per-slot pump (on_poll, the HTTP/WS/SSE inline pump, the file/chunk send pumps).Cross-thread synchronization primitives (no hot-path locks):
_Atomic (acquire/release) on TcpConn.state, rx_head, rx_tail.TcpEvt queue (producer -> owner worker).tcpip_api_call marshaling for the app->lwIP direction (see TX below).Every layer that sends bytes calls the transport API; nobody calls lwIP tcp_* directly:
The same marshal rule covers every raw lwIP call, not just app data: the TCP listener bring-up (TcpListener.add / TcpListener.stop), the UDP transport (protocore_udp_*, used by SNMP / CoAP / captive-DNS / syslog / telemetry), the outbound client (TcpClient), and the DNS resolver all route their tcp_* / udp_* through tcpip_api_call. This is mandatory on arduino-esp32 3.x, where lwIP core-locking asserts on a raw call from any task but tcpip_thread (see docs/BUGS.md); keeping raw lwIP out of the app and worker tasks is the only thing this layer exists to enforce.
Inbound:
Receive-window flow control: now single-owner (transport). recv_cb no longer ACKs on copy. The worker calls ConnPool.ack_consumed once per slot per loop and transport reopens the TCP window by exactly the bytes drained since the last ACK (ack-on-consume; tcp_recved marshaled). The window therefore tracks ring occupancy and a slow consumer cannot overflow the ring. TCP-level requirement: RX_BUF_SIZE >= TCP_WND (you cannot advertise a window larger than your buffer). See docs/BUGS.md "RX flow-control deadlock".
RX ring read API: now single-owner (transport). Consumers no longer index rx_buffer or advance rx_tail. They drain through the transport read API - ConnPool.available / ConnPool.read_byte / ConnPool.read / ConnPool.peek / ConnPool.consume (declared in transport/tcp/protocol/protocol.h, single-consumer per slot). Migrated: presentation.c (HTTP, including the TLS pump), websocket.c, telnet.c, src/network_drivers/presentation/ssh/network/network.c (the SSH socket seam), and the conn_pool-ring services modbus.c / opcua.c (their duplicated ring_peek/consume/avail are now thin adapters over the API). The read functions only consume; the window is reopened by the worker's single ConnPool.ack_consumed per loop - so there is exactly one place that touches the ring indices for draining and one that ACKs.
The device's clients (http_client, mqtt, ws_client) do not each own a raw lwIP stack: they share **TcpClient** (network_drivers/transport/tcp/client/client.*), the client-side peer of the server transport. It is a small fixed pool of outbound connections with the same rules - every raw tcp_*() marshaled to tcpip_thread, a per-connection wire ring, and ack-on-consume (TcpClient.read reopens the window as the caller drains; PROTOCORE_CLIENT_RX_BUF >= TCP_WND), so client and server share one flow-control model. TLS clients layer protocore_tls_client_session_* on top, pointing the BIO at TcpClient.send / TcpClient.read (the ring carries ciphertext).
The rx ring inside mqtt.c / ws_client.c is a separate plaintext frame buffer (post-decrypt for TLS, the assembly buffer the protocol parser reads), owned solely by that module - the client mirror of the server's http body[] vs the conn_pool wire ring. Not cross-layer, correct as-is.
Both transports use ONE shared primitive for the whole ring: ring.h (the _Atomic SPSC index wrapper + the drain math protocore_ring_available / read_byte / read / peek / consume AND the fill math protocore_ring_free / protocore_ring_write_span). The server (ConnPool + recv_cb) and client (protocore_client_* + cc_recv) are thin wrappers over it, so the ring invariants - wrap, ordering, lossless backpressure - live in a single place and no layer reimplements them by hand. Both recv callbacks bulk-memcpy each pbuf span and publish head once. The client ring's indices were volatile; they are now _Atomic like the server (correct cross-core acquire/release ordering).
http_parser_set_stream_hooks(begin, data, abort) are global singletons (last-registered-wins, so OTA / upload / WebDAV streaming are still mutually exclusive per build). All three now take HttpReq*. A sink can keep per-connection state: WebDAV holds per-slot PUT state (s_davput.put[MAX_CONNS] in src/server/io/webdav_handler/webdav_handler.c) and each connection streams to its own file. This fixed the concurrent-PUT clobber (docs/BUGS.md) - HW: 4 parallel PUTs with distinct payloads, all byte-exact.
| Concern | Owner (target) | Status |
|---|---|---|
| Socket TX | transport ConnPool | DONE |
| RX receive window | transport | DONE (ConnPool.ack_consumed, ack-on-consume) |
| RX ring read/drain | transport (read API) | DONE (ConnPool.read*; consumers off the ring) |
| Streaming sink state | per-slot, slot-aware | DONE (s_davput.put[MAX_CONNS], slot-aware hooks) |
| Event routing | session (owner queue) | DONE |
| Working memory | mmgr plain (per slot) | DONE (worker resets it per dispatch) |
| Key material | mmgr secure (per slot) | DONE (disjoint region; reclaiming wipes) |
| Byte moves / parsing | mmgr (mem, str, swar) | DONE (one owner per operation) |
| Outbound client I/O | transport (TcpClient) | DONE (pooled, ack-on-consume; all clients use it) |
ConnPool.available / ConnPool.read_byte / ConnPool.read / ConnPool.peek / ConnPool.consume (transport/tcp/protocol/protocol.h).rx_tail modulo remains. The read functions consume only; ConnPool.ack_consumed stays the only place that reopens the window (per loop), so draining and ACKing each have exactly one owner. HW: 10/10 50 KB byte-exact, backpressure 0.HttpStreamDataCb(HttpReq*, ...) + per-slot WebDAV PUT state s_davput.put[MAX_CONNS]; fixed the concurrent-PUT bug (HW: 4 parallel PUTs, distinct payloads, all byte-exact).TcpClient, the unified client transport; brought to the same ack-on-consume flow control as the server (PROTOCORE_CLIENT_RX_BUF >= TCP_WND). Each module's rx is its own plaintext frame buffer (the client mirror of the server's body[]), module-owned and correct as-is.All phases complete: every cross-layer concern (server TX/RX, RX window, RX read, streaming sink state, events, memory, outbound client I/O) has exactly one owner behind a clean API, and the server and client ring drain math is a single shared primitive (mmgr/ring.h) - two connection pools, one ring/read core.
The data-piping axis above (who owns RX/TX/window/events) is settled. This is the other axis: how each application protocol attaches to that plumbing. The rule is the same - one uniform seam - and it is mostly, but not fully, met.
**The seam - ProtoHandler (server/core/proto_handler.h).** A connection-oriented (TCP) protocol is a vtable of four nullable callbacks keyed by ProtoConn:
dispatch_event() routes each drained TcpEvt to on_{accept,data,close} by conn_pool[slot].proto; handle() calls on_poll for each active slot. Every handler reads its bytes through the transport RX API (ConnPool.read copy-out, or ConnPool.peek+ConnPool.consume zero-copy - never the ring internals) and writes through ConnPool.send/ConnPool.flush. So Telnet, SSH (+ PROTO_SSH_RFWD), Modbus, and OPC UA are fully homogeneous: each is a module that exposes a ProtoHandler and touches the core only through those two APIs.
Connectionless (UDP) services (SNMP, CoAP, DNS, syslog, flow-export) attach through a different but deliberately separate seam - UdpListener.listen (port, handler, handler_ctx), one datagram-in/datagram-out callback. This heterogeneity is correct, not a defect: UDP has no accept/close/slot lifecycle, so folding it into the slot-based ProtoHandler table would be a forced fit. Two transport models, two matched seams.
The request/response core is protocol(version)-agnostic: every version decodes into the shared HttpReq and converges on one match_and_execute / route / Handler (HTTP/1.1, HTTP/2, and HTTP/3), and the response funnels back through the symmetric TX seam - a per-connection protocore_resp_sink_fn function pointer (http_resp_sink[slot], session.h) that HTTP/2 installs at ALPN and HTTP/3 at dispatch. send_text() / send_empty() call http_resp_sink[slot](slot, code, content_type, payload, body_len) when it is set (h2 frames the reply as HEADERS+DATA on the stream; h3 as an HTTP/3 response on its QUIC stream) and otherwise build the HTTP/1.1 message - so the response methods name no protocol. It is the RX ProtoHandler seam's TX counterpart: request decode and response encode both sit behind one uniform per-connection seam.
The fully expanded twin of the simplified chart in the README - the same top-to-bottom waterfall, but every public API method, every registered protocol, and every Layer-6 module on disk is listed.
Generated from the public API,
protocore_builtins.c, andpresentation/bytools/ci_tooling/generate/gen_api_flow.py- do not edit by hand. This is the fully expanded twin of the simplified request-lifecycle chart in the README: the same top-to-bottom waterfall, but every public method, every registered protocol, and every Layer-6 module on disk is listed (nothing is capped). Color is the OSI layer; the green path is the response. Mermaid source:diagrams/api_flow_detail.mmd.
session.c (L5) is now protocol-agnostic - DONE.** The dispatcher owns only the mechanism (register / look up / route / drain) and names no protocol. Each protocol's handler lives in its own module and is exposed by a pure accessor (HttpConn.proto_handler in presentation, ssh_protocore_handler() in ssh/server, ...) that carries no dependency on the session layer. The single policy file protocore_builtins.c (src/server/) maps each built-in to its accessor behind its feature flag; proto_get() calls protocore_register_builtins() once, lazily, so the native harness still works before proto_begin(). Adding a protocol = write its module + one guarded line in protocore_builtins.cprotocore_ssh_forward_begin().)presentation.c (Layer 6, already the HTTP-connection glue) now owns the HTTP ProtoHandler: http_evt_{accept, data,close} plus tls_data (the TLS handshake pump + ALPN "h2" detection + WebSocket upgrade check before the HTTP/1.1 parser). L5 no longer includes TLS / http2 / websocket / http_parser.on_poll seam - DONE. HTTP's poll (the file/chunk send pumps, the WebSocket + SSE drains, the keep-alive re-parse, request dispatch) is instance-bound - it dispatches into the route table - so it used to be a large inline block in the worker loop guarded by if (proto != PROTO_HTTP). That block is now protocore_http_on_poll(), installed as the HTTP ProtoHandler's on_poll (via HttpConn.set_poll, the on_poll analogue of the http_resp_sink TX seam; it is wired in at the top of service_once()). The worker dispatch loop now calls on_poll uniformly for every protocol including HTTP - it names no protocol and has no special case. The singleton pollers (ssh, rfwd) gate on CONN_ACTIVE inside their own on_poll, preserving the behavior the loop-level gate used to give them.send_text() / send_empty() no longer branch on conn->h2 / conn->h3; a self-framing protocol installs a http_resp_sink[slot] function pointer (h2 at ALPN, h3 at dispatch) and the response methods route through it, so the L7 responders name no protocol. This is the TX counterpart of the RX ProtoHandler seam; adding a self-framing protocol means installing one http_resp_sink, not editing the responders.Net: L5 is pure dispatch and every protocol (including HTTP) lives behind the same uniform seam via its own module - request decode through the ProtoHandler seam (accept / data / close / poll), response encode through the http_resp_sink seam. The worker dispatch loop names no protocol and has no special case: HTTP plugs in exactly like SSH, Telnet, Modbus, or OPC UA. The only remaining inherent trait is that TLS is an HTTP-only inline transform (item 5). The piping is straight.
One module owns every byte the library hands out and every operation that moves one. Nothing under src/ declares its own working storage. A caller borrows from a pool, and the pool is the only thing that knows where the bytes are, how many there are, and who may reach them.
This is the general rule, seen from memory: a function that needs a resource requests it from that resource's owner. It does not reach past the owner, it does not keep a private copy, and it does not decide for itself where the bytes live. For memory the owner is mmgr, the answer is one of exactly two pools, and the request is a call.
protocore_arena is the mechanism and there is exactly one of it: one contiguous region with two allocators growing toward each other and the free space floating in the middle.
The persistent end is a first-fit free list - individual free in any order, adjacent-block coalesce, top-block shrink, and the bytes come back zeroed. The scratch end is a bump with O(1) reset and mark/release savepoints, aligned up to PROTOCORE_ARENA_MAX_ALIGN (16). Whichever side needs room takes it, and both ends fail closed (NULL) before they can cross. protocore_arena_set chains a DRAM base and a PSRAM extension: a borrow takes the first region that fits, a free routes to the owning region by address. No heap, no stdlib, all state in protocore_arena (no globals), so it is unit-tested on the host.
plaintext.c and secure.c are not second allocators. Each is an access layer that instantiates that one mechanism over its own compile-time-sized storage and decides who may reach it. Nothing outside src/mmgr/ names protocore_arena.
plaintext (plain) | secure (secure) | |
|---|---|---|
| holds | anything whose bytes are not secret | key material only |
| hand-out | uninitialized | uninitialized |
| reclaim | move the offset | wipe, then move the offset |
| who reclaims | the worker, once per dispatch | the borrower, at its own mark/release |
| long-lived storage | (none - the pool is emptied each dispatch) | secure.persist_span(), the end no mark walks |
| per-slot size | PROTOCORE_PLAINTEXT_ARENA_SIZE | PROTOCORE_SECURE_ARENA_SIZE, derived from the enabled crypto |
The pool is chosen by whether the contents are secret and by nothing else. Lifetime is not the axis: both pools carry long-lived and ephemeral borrows. A peer's public point, a ciphertext on its way out, a staging buffer for an outbound frame are plaintext; putting them in the secure pool only shrinks the room left for real secrets.
Each pool is cut one arena per slot: one per server worker (PROTOCORE_WORKER_COUNT), plus the ghost at PROTOCORE_GHOST_WORKER_SLOT, the library's own. A borrow resolves its slot from protocore_worker_self() and never takes one as an argument. A borrow cannot cross workers and the bump needs no lock. A caller that is not a server worker clamps to the ghost. It does not fall back to worker 0. Under PROTOCORE_DEBUG_CHECKS each slot records the first execution context to touch it and asserts on a second, turning a future cross-core mistake into an immediate visible failure.
That identity lives in arena.h, with the thing it indexes: mmgr arbitrates the memory, so it owns the identity the arbitration keys on. Starting, waking and stopping those workers is scheduling and stays in server/core/worker.h.
dispatch_event() calls protocore_plaintext_reset() before handing an event to its protocol handler. A borrow is valid only until that handler returns and a forgotten release cannot accumulate across events. Inside one dispatch, plain.mark() / plain.release() nest.protocore_secure_release() zeroes the reclaimed extent before the position moves, so the bytes are already zero at the instant they become available again - there is no window in which the next borrow is handed the previous tenant's key material. The wipe is structural. A discipline every return path has to remember had already been missed on two SSH key-exchange error paths. protocore_secure_wipe() is the same primitive for storage that was never in a pool.secure.persist_span(n) takes the end no mark walks. A credential table or a key schedule bound once at setup survives every release and every reset, and comes back zeroed.The two pools are disjoint regions, so the owner is recoverable from the pointer alone. protocore_plaintext_owns() and protocore_secure_owns() are one unsigned subtract and compare against a compile-time extent, with no loop, no per-slot comparison and no per-allocation metadata, and they are mutually exclusive by construction: a secret can never be accepted where plaintext is expected, or the reverse. A pointer below the base wraps to a huge offset and fails the same bound as one past the end. An overrun cannot test as still-inside. slot_of() answers which slot. A borrow being handed back can be asserted to belong to the calling worker.
high_water() reports the peak any slot reached - the number to size the arena by. Each pool's backing storage is named only from its bind(), on the allocation path, so --gc-sections reclaims the whole block from firmware that never borrows; that is why reset() must not bind.
A pool size is not a chosen number. Each translation unit states the worst-case bytes it borrows in a single call as a PROTOCORE_WORK_* constant in protocore_config.h, and the module that owns the working set proves it where the struct lives:
PROTOCORE_SECURE_ARENA_SIZE is then the sum of the terms a build actually compiles, each gated by its feature flag. A build pays only for the code it has. The sum is a strict upper bound however those working sets nest. A deepest-nest figure is only correct while the call graph stays as it is. It buys certainty with a little slack.
The consequence follows. A module whose working set grows past its declaration fails the build, naming itself, instead of exhausting the pool at run time on a part nobody was watching. These are sizes and not offsets, so nothing couples one module to another - order is irrelevant and adding a module shifts no one.
The RX rings are the exception, and they are the only one: TcpConn::rx_buffer[] in the static conn_pool, and the UDP datagram rings. They exist because the producer is not a worker - tcpip_thread fills them - so they cannot be a worker-slot borrow. The owning worker drains its ring into a pool borrow, and everything from there on is a pool pointer:
TX is the mirror: the send functions read out of a plaintext or secure borrow on the owning worker and write to the wire. A byte is never staged into a third buffer on the way out.
Borrowing and moving live in the same module because both are about memory the library owns. Every operation below acts on a pool borrow, and each is stated once. A fix lands in one place and every codec inherits it.
protocore_span / protocore_cspan** (span.h) carry storage, capacity, the produced length, and a sticky overflow flag. pos keeps counting past cap on overflow. An undersized region reports the capacity it should have had instead of only failing. A failed borrow yields {NULL, 0}, never a null with a live capacity. A caller that skips protocore_span_ok() writes nothing and never dereferences null. plain.span() / secure.span() are the preferred borrow: one argument sets both fields, so the length cannot drift from what was reserved.mem** (protomem) walks a span a register word at a time; a source not co-aligned with the destination is funneled through two shifts and an OR. mem.cmp is not constant time - a secret comparison uses protocore_ct_eq (crypto/ct_eq.h).str** (protostr) answers where a bounded run ends and where two part company, one word per test, with ci folding ASCII case inside one body. Also the no-stdlib number parsing (to_long / to_ulong / to_double / to_float).swar** is the access layer under both: load a word, test its lanes branchless, name the lane that fired. Byte order enters in exactly one place, protocore_swar_zero_lane. Nothing in it walks a buffer or takes a capacity. That keeps the claim true. The walks built on it are str: shared/runops.h was a second full implementation of the same operations and was removed on 2026-08-08 (docs/BUGS.md), with its 44 call sites rewritten onto str.rawmemcpy.h** owns the load itself: PROTO_RAW (aligned(1) + may_alias) for an address that carries no guarantee, proto_al_load for one the caller has proven, and proto_raw_read for a span move at any alignment. PROTO_RAW_WORD follows PROTO_WORD_BITS from the board profile, not the build machine's pointer width.membuild / protoframe** build into bytes the caller already owns and never allocate. protocore_sb latches ok false on the first append that does not fit, so the caller tests one flag at the end instead of a return value per call. A frame is a static const protocore_field[] in rodata walked by one engine, so nothing is parsed at runtime and no float formatter is linked unless a frame declares a float field.ring.h** holds the SPSC drain and fill math both transports share (see the RX path above), plus the segment ring and the slot-mask view.The library is already OSI-layered and the code prefix (protocore_ / PROTOCORE_) is vendor-neutral, but three layers still bake in ESP silicon: the per-die board profiles, the crypto accelerator HAL, and the physical (EMAC + PHY + raw-register / lwIP glue) layer. The goal is to move the vendor-neutral majority into common areas and let the preprocessor pull in exactly one vendor backend per build, so adding STM32 (then TI Sitara / CC32xx, RP2350, others) is "add a subdir + wire the
selector", not a fork. We are targeting a broad board matrix; the seams have to be designed once, correctly, up front.
Layout - common by default, a thin per-vendor subdir for what is genuinely silicon-specific:
Selector (the only new common seam): a single protocore_platform.h maps the toolchain's target macro onto two axes and nothing else pulls vendor detail directly:
PROTOCORE_VENDOR_ESP (from CONFIG_IDF_TARGET_*), PROTOCORE_VENDOR_STM (from STM32* / CMSIS device), PROTOCORE_VENDOR_RP (PICO_RP2350 ...), PROTOCORE_VENDOR_TI.CONFIG_IDF_TARGET_*-style discriminator per vendor.A common selector point then resolves the backend once per layer: #if PROTOCORE_VENDOR_ESP -> #include "esp/..." #elif PROTOCORE_VENDOR_STM -> stm/... else -> the portable software path (this is how board_profile.h picks the die profile). The crypto layer already does the vendor-agnostic thing without a dispatcher: crypto/ is portable C, and each TU keys off the HAL's capability macro (PROTOCORE_RSA_MODMUL_HW) that the selected test/core_setup/hal/<vendor>/ backend defines - so crypto/ never moves and never names a vendor. Common code sees an API/macro, not a vendor subdir.
Principles (carry the ones the ESP crypto HAL already proved):
crypto/. A brand-new vendor with no accelerator still links and runs from day one; accel is added incrementally.PROTOCORE_ register map, no HAL_* / esp_* / vendor struct (the esp_crypto_hal rule, applied per vendor). STM32 backends poke CRYP/HASH/PKA registers directly.static_assert regmap cross-check (penetration_testing/rig_firmware/hal_verify today for ESP soc macros; add an STM CMSIS variant). A map is proven correct even for silicon we have no board for, plus an on-device KAT where a board exists.server/clock/clock.h time) behind a thin services/protocore_rtos. A vendor picks its RTOS without touching callers.stdlib in src/.Sub-items (sized):
esp/ subdirs + add the protocore_platform.h selector, with zero behavior change on ESP32 (pure move + include rewire, CI-gated). (M)**Rename - drop "esp" from the name (DeterministicAsyncWebServer).** The code prefix protocore_/PROTOCORE_ is already vendor-neutral ("Deterministic Web Server"), so this is a product/library-name + docs change, not a code-wide symbol churn. It touches the repo/library display name (library.json / library.properties), README, and the doc prose that says "ESP" where it now means "any target". Do it alongside the STM32 backend landing (a real second vendor), not before, so the name stops being aspirational the moment it changes. (S, coordinated)
Found while benching the connected rigs (2026-07-26). Big-integer modexp (DH-2048, RSA sign/verify) already routes through mbedtls with CONFIG_MBEDTLS_HARDWARE_MPI=y on every die - there is no software fallback being wrongly taken on firmware. But the mbedtls path uses CONFIG_MBEDTLS_MPI_USE_INTERRUPT=y: it blocks on the RSA-done interrupt for every modular multiply, and that per-op round-trip, not the accelerator, dominates. Evidence: our own polling-mode HAL fe_mul (single-shot protocore_rsa_modmul) is comparable across dies (S3 1403 / P4 1896 / C6 1695 cyc for a 256-bit MODMULT), yet mbedtls's 2048-bit DH modexp spreads ~7.5x (P4 ~20.7k cyc per 2048-bit modmul vs a raw modmul that should cost only a few thousand). The overhead is the interrupt/driver layer, not the silicon.
Opportunity: build a PC modexp on top of the crypto HAL's polling protocore_rsa_modmul (Montgomery, CRT for RSA sign, constant-time exponent handling) so RSA/DH run at the accelerator's real throughput instead of interrupt-round-trip-bound. This is also a portability win: it gives the library a vendor-agnostic HW modexp (the same HAL API the STM PKA / others implement), so RSA/DH stop depending on each vendor's mbedtls port. Tradeoff to measure, not assume (run the experiment): polling busy-waits the worker for the modexp duration where interrupt mode yields; for the deterministic single-owner model a bounded ~tens-of-ms blocking op is likely fine, but confirm against PROTOCORE_WORKER_COUNT scheduling before switching the default. Keep mbedtls as the fallback where the HAL has no MODMULT (classic ESP32) or where a die's interrupt path already wins (measure C6). Legacy finite-field DH is lower-priority than RSA sign; modern KEX is curve25519/ECDH already.
Raised 2026-07-27. lwIP is a fine portable reference stack but it is slow in the places that matter for a deterministic single-owner server, and several of its costs are structural, not tunable. Mirror what the crypto HAL did (test/core_setup/hal/: direct registers, our own PROTOCORE_ register map, zero soc/ / vendor symbols, ground-truth static_assert-verified vs the vendor headers): pull the networking data path out of lwIP into a direct-register/DMA HAL under network_drivers/physical/<vendor>/, and keep lwIP only as the portable fallback / cold-path L3+ where throughput does not matter.
lwIP pain points (measure each, fix the ones that pay):
tcpip_thread marshaling** - every socket/PCB call is marshaled onto the single tcpip task via tcpip_api_call (our transport layer already pays this - real per-op overhead + a serialization point); a direct datapath skips the trip.tcp_pcb in TIME_WAIT for ~1 min (observed 2026-07-27 on the SMB rig: free heap dipped exactly one PCB per rapid outbound probe, then recovered on expiry). A lean connection table with our own reuse policy reclaims it immediately for the trusted deterministic case.Opportunity: a protocore_net datapath HAL (RX/TX descriptor rings + a minimal, deterministic TCP fast-path) that runs the common case at wire/DMA speed, with lwIP retained for ARP/ICMP/DHCP/edge cases and as the portability floor for a new vendor before its HAL exists. Same shape as the crypto HAL: vendor-agnostic API, per-die register backends, ground-truth verified vs the vendor headers.
Tradeoffs to measure, not assume (run the experiment): a hand-rolled TCP fast-path must stay RFC-correct - re-verify interop against real peers, not just self-consistency (a self-consistent stack test proves only self-consistency). Scope per pain point; land the zero-copy DMA ring + TIME_WAIT reclaim first (highest value, lowest correctness risk), and defer a full TCP fast-path until measured. Keep every change behind the network_drivers/physical/<vendor>/ seam so lwIP-only vendors still boot.
The multi-vendor section above describes where silicon-specific code goes. This section records the harder questions that surfaced once we started actually going wide, and what was decided. Targets: Xtensa, RISC-V, Arm, TI C2000.
1. uint8_t is not an octet. This is the deepest portability problem in the library and it is not a naming issue. TI C28x has no 8-bit addressable memory: CHAR_BIT == 16, so uint8_t is 16 bits wide and sizeof counts 16-bit words. ProtoCore is essentially one large wire-protocol byte machine - every uint8_t buf[], every memcpy into a frame, every length in octets, every crypto block, every sizeof used as a wire length. An array of 100 uint8_t is 100 sixteen-bit words there, and a 1500-byte frame does not lay out the way the code assumes. C-versus-C++ is a footnote next to this.
2. Control law is reviewed as C. On C2000 parts the control-law code is written and reviewed as C. An API that requires a C++ compiler to call is unusable exactly where the library most needs to be usable.
3. Vendor idioms leak upward. The ESP-IDF component build currently hard-requires arduino-esp32 because the core calls Arduino APIs (WiFi, ESP, millis, Serial). Every such call is a porting blocker for every other architecture.
4. Guarantees have to survive a compiler. "Deterministic" and "constant-time" are claims about emitted instructions. Asserting them in a comment does not make them survive a toolchain upgrade.
The API and the implementation are both C11. Flat protocore_ / PROTOCORE_ names at global scope, no namespace. A C caller can reach everything (see SYMBOLS.md for the full naming law and the designs rejected). src/ carries no .cpp at all; the three vendor-wrapper exceptions under test/core_setup/ are listed in SYMBOLS.md.
One octet abstraction, packed everywhere.
A native C28x pointer cannot name an octet, but a house-defined one can. Zero cost on Xtensa / RISC-V / Arm, where the whole abstraction vanishes at compile time. The unpacked one-octet-per-word variant (2x RAM) was considered and dropped - it is not needed, because the cost of packing is not pervasive:
hdr[3], fixed-layout header and framing codecs - most of the wire code): the lane is a compile-time constant, the mask and shift fold away, and it becomes a word load plus an immediate AND. Melts entirely.base + CONSTANT with no runtime index. That leaves one bit of uncertainty per record, removable by requiring word-aligned record bases or dispatching on base parity once at record entry.Still to pin: which lane is octet 0 (must be invariant across targets, not the machine's word endianness, or the same frame serializes differently per platform); the unwrap-to-native operation for DMA/MAC handoff and its alignment precondition; and the fact that sizeof stops being a wire length, so octet-count expressions must replace it - which is mechanically checkable.
No vendor language or idioms in the core. Vendor registers reach the library only through a HAL that auto-configures per variant capability, never through a chip check. This extends the pattern already shipped in test/core_setup/hal/ (direct registers, house-owned register map, zero vendor symbols) to every subsystem, and it is what has to happen before the build system can stop depending on arduino-esp32.
Guarantees are proven at the binary. Where the library promises a behavior, the promise is checked against emitted instructions and documented as claim -> disassembly -> why. The claims that get this treatment: constant-time crypto (no branch or memory access depends on a secret), no-heap-after-begin() (no allocator reachable in the relevant .text), and bounded ISR / critical-section paths (a counted worst case, not an estimate). The octet abstraction earns a fourth once it lands: that it compiles away on byte-addressable targets and strip-mines to word moves on C28x. Each is measured.