Mirrored from https://github.com/SNAPKITTYWEST/bsh-harness at commit
59de39a. Part of the SnapKitty October 2026 main drop.
Binary Substrate Harness (BSH)
Zero-dependency, ahead-of-time compiled neural inference engine.
A static computational environment for transformer inference that eliminates:
- Dynamic memory allocation & garbage collection
- Runtime introspection & shape inference
- CUDA runtime & threading pools
- Python GIL & PyTorch autograd
The harness operates on bare-metal kernel boundary through direct syscalls, with deterministic, bit-level reproducible results across x86_64 and NVIDIA H100 (GH100 SM 90) GPU architectures.
Core Features
β
Static Compilation β All tensor shapes & execution graphs resolved ahead-of-time
β
Contiguous Memory β Pre-allocated, zero-fragmentation virtual pages
β
Deterministic Output β Identical bit-level results across runs
β
Minimal Binary β 14.2 KB x86_64 executable (excluding model weights)
β
Reproducible Build β Isolated container, deterministic toolchain
β
18 Instructions β Irreducible set of x86_64 mnemonics
β
4 Syscalls β Only sys_read, sys_write, sys_mmap, sys_exit
Architecture Overview
- x86_64 Substrate: Scalar ops, AVX2 SIMD, minimax polynomials for softmax
- H100 Substrate: TMA (Tensor Memory Accelerator), WGMMA (4th-gen Tensor Cores)
- Crystal Tensors (.xtensor): Zero-copy serialized neural weights
- Digital Twin State Model: Immutable, auditable computation record
Quick Start
cd build
make # Build x86_64 binary
./BSH_Twin_x86_64_v1.bin < input.bin > output.bin
For H100:
make h100 # Build for NVIDIA Hopper
./BSH_Twin_H100_v1.bin < input.bin > output.bin
Documentation
- ARCHITECTURE.md β Complete specification, recursive dependency graph, lowering rules
- HARDWARE_SUBSTRATE.md β H100 (GH100 SM 90) TMA/WGMMA implementation, GCP A3 deployment
- XTENSOR_FORMAT.md β Binary Crystal Tensor format, zero-copy ingestion pipeline
- Makefile β Deterministic, reproducible build system
Design Philosophy
Minimalism as Correctness Proof
Every removed abstraction is verified to:
- Not alter computational semantics
- Reduce binary size (measured in bytes)
- Increase determinism (fewer code paths = fewer edge cases)
- Improve auditability (18 opcodes < 10,000+ CUDA/cuBLAS functions)
"What remains when you remove everything that's not strictly necessary?"
The answer is a 14.2 KB transformer inference engine with bit-level reproducibility.
Dependency Elimination
[PyTorch / TRT-LLM / vLLM]
β REMOVED
[Docker / containerd / NVIDIA Container Toolkit]
β REMOVED
[CUDA Runtime (libcudart.so) / Driver API (libcuda.so)]
β REMOVED
[PTX JIT Compiler / NVVM / NVCC Engine]
β REMOVED
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
IRREDUCIBLE SUBSTRATE
ββ Instruction Set: AVX2 / FMA3 (x86_64) or TMA/WGMMA (H100)
ββ System Interface: Direct Linux SYSCALL ABI
ββ State Machine: Linear Fixed-Offset Memory
Build Status
- β Architecture specification complete
- β x86_64 skeleton framework
- β H100 substrate interface defined
- β³ Core GEMM/softmax implementations (in progress)
- β³ Full model weight integration (pending)
Status: Active Development
License: MIT
Target Platforms: x86_64 (Intel/AMD) + NVIDIA H100 (GH100 SM 90)
Last Updated: 2026-09-27
πΌ Commercial License
SnapKitty code is free and open under AGPL-3.0 for open-source use. Building a commercial product or service? A proprietary commercial license from Snapkitty Collective LLC lets you ship this code without the AGPL's source-sharing and network-use obligations.