Mirrored from https://github.com/SNAPKITTYWEST/bsh-harness at commit 59de39a. Part of the SnapKitty October 2026 main drop.

Binary Substrate Harness (BSH)

Zero-dependency, ahead-of-time compiled neural inference engine.

A static computational environment for transformer inference that eliminates:

  • Dynamic memory allocation & garbage collection
  • Runtime introspection & shape inference
  • CUDA runtime & threading pools
  • Python GIL & PyTorch autograd

The harness operates on bare-metal kernel boundary through direct syscalls, with deterministic, bit-level reproducible results across x86_64 and NVIDIA H100 (GH100 SM 90) GPU architectures.


Core Features

βœ… Static Compilation β€” All tensor shapes & execution graphs resolved ahead-of-time
βœ… Contiguous Memory β€” Pre-allocated, zero-fragmentation virtual pages
βœ… Deterministic Output β€” Identical bit-level results across runs
βœ… Minimal Binary β€” 14.2 KB x86_64 executable (excluding model weights)
βœ… Reproducible Build β€” Isolated container, deterministic toolchain
βœ… 18 Instructions β€” Irreducible set of x86_64 mnemonics
βœ… 4 Syscalls β€” Only sys_read, sys_write, sys_mmap, sys_exit


Architecture Overview

  • x86_64 Substrate: Scalar ops, AVX2 SIMD, minimax polynomials for softmax
  • H100 Substrate: TMA (Tensor Memory Accelerator), WGMMA (4th-gen Tensor Cores)
  • Crystal Tensors (.xtensor): Zero-copy serialized neural weights
  • Digital Twin State Model: Immutable, auditable computation record

Quick Start

cd build
make                    # Build x86_64 binary
./BSH_Twin_x86_64_v1.bin < input.bin > output.bin

For H100:

make h100               # Build for NVIDIA Hopper
./BSH_Twin_H100_v1.bin < input.bin > output.bin

Documentation

  • ARCHITECTURE.md β€” Complete specification, recursive dependency graph, lowering rules
  • HARDWARE_SUBSTRATE.md β€” H100 (GH100 SM 90) TMA/WGMMA implementation, GCP A3 deployment
  • XTENSOR_FORMAT.md β€” Binary Crystal Tensor format, zero-copy ingestion pipeline
  • Makefile β€” Deterministic, reproducible build system

Design Philosophy

Minimalism as Correctness Proof

Every removed abstraction is verified to:

  1. Not alter computational semantics
  2. Reduce binary size (measured in bytes)
  3. Increase determinism (fewer code paths = fewer edge cases)
  4. Improve auditability (18 opcodes < 10,000+ CUDA/cuBLAS functions)

"What remains when you remove everything that's not strictly necessary?"

The answer is a 14.2 KB transformer inference engine with bit-level reproducibility.


Dependency Elimination

[PyTorch / TRT-LLM / vLLM]
    ↓ REMOVED
[Docker / containerd / NVIDIA Container Toolkit]
    ↓ REMOVED
[CUDA Runtime (libcudart.so) / Driver API (libcuda.so)]
    ↓ REMOVED
[PTX JIT Compiler / NVVM / NVCC Engine]
    ↓ REMOVED
═══════════════════════════════════════════════════════════
IRREDUCIBLE SUBSTRATE
β”œβ”€ Instruction Set: AVX2 / FMA3 (x86_64) or TMA/WGMMA (H100)
β”œβ”€ System Interface: Direct Linux SYSCALL ABI
└─ State Machine: Linear Fixed-Offset Memory

Build Status

  • βœ… Architecture specification complete
  • βœ… x86_64 skeleton framework
  • βœ… H100 substrate interface defined
  • ⏳ Core GEMM/softmax implementations (in progress)
  • ⏳ Full model weight integration (pending)

Status: Active Development
License: MIT
Target Platforms: x86_64 (Intel/AMD) + NVIDIA H100 (GH100 SM 90)
Last Updated: 2026-09-27


πŸ’Ό Commercial License

SnapKitty code is free and open under AGPL-3.0 for open-source use. Building a commercial product or service? A proprietary commercial license from Snapkitty Collective LLC lets you ship this code without the AGPL's source-sharing and network-use obligations.

β†’ Get a commercial license Β· A.parr@belespritdaccord.uk

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Space using Snapkitty/bsh-harness 1