47 systems · 5 layers
The modern AI stack, rebuilt to understand it
Every row is a standalone repository. Most are study builds; the nineteen marked ⚙ are engineered libraries with tests and strict typing. These are the stars in the console's sky — the shipped, production work is on the work page.
| System | Layer | What it is | Lang | Repository link |
|---|---|---|---|---|
| Transformer | Modeling & training | “Attention Is All You Need” implemented end-to-end. | Python | src ↗ |
| Vision Transformer⚙ | Modeling & training | Production-grade ViT: training engine, FastAPI serving, ONNX export, >95% test coverage. | Python | src ↗ |
| CLIP multimodal projector⚙ | Modeling & training | From-scratch dual encoders + symmetric contrastive loss, packaged as a hardened FastAPI embedding service with OpenAPI and operational guides. | Python | src ↗ |
| WhisperLite⚙ | Modeling & training | Whisper-style ASR from scratch: log-mel frontend, byte BPE, encoder–decoder Transformer, KV-cached decoding, training and hardened serving — 216 tests. | Python | src ↗ |
| Diffusion models | Modeling & training | DDPM + DDIM with a timestep-conditioned UNet — every equation implemented, no `diffusers`. | Python | src ↗ |
| Small language model | Modeling & training | A small LM trained from scratch, end to end. | Python | src ↗ |
| State space model | Modeling & training | Mamba-style selective state space model. | Python | src ↗ |
| MoE router | Modeling & training | Mixture-of-Experts routing layer with load-balancing losses. | Python | src ↗ |
| Knowledge distillation | Modeling & training | Teacher→student distillation pipeline. | Python | src ↗ |
| BPE tokenizer | Modeling & training | Byte-pair encoding tokenizer from first principles. | Python | src ↗ |
| FlashAttention | Modeling & training | Tiled, IO-aware attention kernel reimplementation. | Python | src ↗ |
| Distributed training | Modeling & training | FSDP / tensor-parallel training loop mechanics. | Python | src ↗ |
| RLHF pipeline | Alignment & fine-tuning | Reward modeling + PPO policy optimization loop. | Python | src ↗ |
| DPO | Alignment & fine-tuning | Direct Preference Optimization loss and trainer. | Python | src ↗ |
| LoRA trainer | Alignment & fine-tuning | Low-rank adaptation training from scratch. | Python | src ↗ |
| PEFT library⚙ | Alignment & fine-tuning | LoRA, DoRA, VeRA, IA³, adapters, prefix & prompt tuning — mypy-strict, ~96% coverage. | Python | src ↗ |
| Inference server⚙ | Inference & serving | ONNX Runtime + Axum: backpressure, graceful shutdown, Prometheus metrics. | Rust | src ↗ |
| Paged KV cache⚙ | Inference & serving | vLLM-style page tables, copy-on-write forking, radix prefix cache, continuous batching. | Python | src ↗ |
| Speculative decoding | Inference & serving | Draft-and-verify decoding with acceptance sampling. | Python | src ↗ |
| Logit processor | Inference & serving | Composable sampling: temperature, top-k/p, repetition penalties. | Python | src ↗ |
| Prompt cache | Inference & serving | Prefix-aware prompt caching for LLM serving. | Python | src ↗ |
| Quantization library⚙ | Inference & serving | INT8 / FP4 (E2M1) / NF4 with MinMax, percentile, KL & MSE calibration — zero `bitsandbytes`. | Python | src ↗ |
| AI gateway | Inference & serving | Multi-provider LLM gateway: routing, retries, keys. | TypeScript | src ↗ |
| funcflow⚙ | Inference & serving | Model-agnostic function router with validated tool registries, dependency-aware plans, parallel execution, retries and provider fallbacks. | TypeScript | src ↗ |
| Vector database | Retrieval & data | HNSW index from scratch — recall >0.95 on a 10K-vector benchmark. | Python | src ↗ |
| RAG pipeline | Retrieval & data | Chunking, embedding, retrieval and generation without frameworks. | Python | src ↗ |
| Graph RAG | Retrieval & data | Knowledge-graph-backed retrieval augmented generation. | Python | src ↗ |
| Semantic router | Retrieval & data | Embedding-based intent routing for LLM apps. | Python | src ↗ |
| Data curation pipeline | Retrieval & data | Dedup, filtering and quality scoring for training corpora. | Python | src ↗ |
| Synthetic data | Retrieval & data | Synthetic dataset generation toolkit. | JavaScript | src ↗ |
| Feature store | Retrieval & data | Offline/online feature storage with point-in-time correctness. | Python | src ↗ |
| Code interpreter | Retrieval & data | Sandboxed code-execution tool for agents. | Python | src ↗ |
| LLM eval harness⚙ | Evaluation & safety | JSONL datasets → prompt templates → provider adapters → scorers → crash-safe artifacts. | Python | src ↗ |
| Guardrails | Evaluation & safety | Input/output validation and policy enforcement for LLM apps. | TypeScript | src ↗ |
| SAE interpretability toolkit⚙ | Evaluation & safety | Activation hooks → sharded SafeTensors → ReLU/Top-K sparse autoencoders → evaluation, feature inspection and causal interventions. | Python | src ↗ |
| ThinkAct agent runtime⚙ | Evaluation & safety | Explicit ReAct state machine with strict JSON actions, bounded memory, tool timeouts, repeat detection and prompt-injection redaction. | Python | src ↗ |
| Reasoning orchestrator⚙ | Evaluation & safety | Decompose → reason → critique → synthesize pipeline with targeted revision cycles, confidence signaling and runtime budgets. | Python | src ↗ |
| AI observability platform⚙ | Evaluation & safety | OTLP tracing, retrieval/agent-trajectory visualization and cost accounting for LLM apps. | Python | src ↗ |
| Durable coding and research agent | Evaluation & safety | Local-first agent: SQL-backed task graph, checkpoint/resume, evidence-cited reports. | Python | src ↗ |
| Text embedding model | Modeling & training | Transformer text embedding model: contrastive training, export, FAISS index and serving. | Python | src ↗ |
| Hybrid distributed training⚙ | Modeling & training | DDP/FSDP/tensor/sequence parallelism from raw collectives, checked against PyTorch DDP. | Python | src ↗ |
| Knowledge graph builder | Retrieval & data | Documents to an evidence-linked graph: ontology validation, entity resolution, retraction. | Python | src ↗ |
| LLM evaluation platform⚙ | Evaluation & safety | Typed monorepo for offline LLM/RAG/agent evals: metrics, paired statistics, gates. | Python | src ↗ |
| Model merger⚙ | Alignment & fine-tuning | Model soups and SLERP checkpoint merging with bounded, streaming, per-tensor memory use. | Python | src ↗ |
| Two-tower recommender⚙ | Retrieval & data | Two-tower candidate retrieval: in-batch softmax, FAISS HNSW, policy-aware reranking. | Python | src ↗ |
| Text-to-SQL engine⚙ | Retrieval & data | NL to SQL with AST validation, tenant rewriting, cost analysis after generation. | Python | src ↗ |
| Text-to-speech pipeline⚙ | Modeling & training | FastSpeech2 + HiFi-GAN vocoder: data pipeline through serving; ships no trained voices. | Python | src ↗ |
The modern AI stack, rebuilt to understand it — 47 readable reference implementations. ⚙ marks the 19 that are engineered libraries (tests, strict typing, production concerns); the rest are study builds. Part of 195 public repos — the applied, shipped work is a separate query (“show shipped systems”).