Skip to content

feat: add Rust CLI with tokenizer models - #2

Merged
eordano merged 3 commits into
mainfrom
rust-cli
Feb 28, 2026
Merged

eordano merged 3 commits into
mainfrom
rust-cli

Conversation

@eordano

@eordano eordano commented Feb 28, 2026

Copy link
Copy Markdown
Owner

Three tokenizer backends implemented from scratch (no tokenizer
libraries): double-array trie for Claude, FNV-1a frozen hash tables
for OpenAI tiktoken (o200k_base), and BPE with frozen hash tables
for 7 HuggingFace models. All model data is embedded at build time
via include_bytes! for zero-copy runtime access.

Remove section comments, deduplicate CSS rules, compress DOM refs
with a $ helper, consolidate event handlers, simplify diff rendering,
and clean up build scripts. Bump version to 0.99.0 and add --version
flag to Node.js CLI.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
@github-actions

github-actions Bot commented Feb 28, 2026 •

Copy link
Copy Markdown
PR Preview Action v1.8.1
Preview removed because the pull request was closed.
2026-02-28 21:11 UTC

Three tokenizer backends implemented from scratch (no tokenizer
libraries): double-array trie for Claude, FNV-1a frozen hash tables
for OpenAI tiktoken (o200k_base), and BPE with frozen hash tables
for 7 HuggingFace models. All model data is embedded at build time
via include_bytes! for zero-copy runtime access.

Includes build.rs for compile-time data generation, flake.nix
integration with rustModelsDir, and updated documentation.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Tiktoken's byte-level BPE was O(n²): linear scan for the best merge
pair + Vec::remove on every iteration. An adversarial input like
"abcdefghijklmnopqrstuvwxyz"×20K (matched as a single regex piece)
took 85s at 500K bytes, scaling quadratically.

Replace with the same priority-queue + linked-list skip structure used
by the HF BPE path: BinaryHeap<(rank, index, generation)> for O(log n)
pops, next/prev/alive arrays for O(1) neighbor traversal, generation
counters for cheap stale-entry invalidation with rank re-verification
as a correctness backstop. Expected ~31x speedup at 500K based on
the HF BPE path benchmarks under the same adversarial load.

Add frozen_map_get_concat() and fnv_hash_concat() to frozen.rs —
zero-allocation lookup of two concatenated slices in a frozen hash
table, avoiding the temporary Vec that the old bpe_count needed for
every pair probe. Remove the now-unused frozen_map_get().

Three other wins stacked in:

  Frozen hash tables: drop the power-of-2 slot count requirement.
  Replace bitmask indexing with Lemire fast range reduction
  (128-bit multiply + shift, no division). Slot counts shrink from
  next_power_of_two(⌈4n/3⌉) to exactly ⌈4n/3⌉ — ~35% fewer slots
  for large models, directly reducing embedded binary size since each
  slot is 14–18 bytes × ~200K entries.

  Claude double-array trie: reference TRIE_BIN as &'static [u8]
  slices instead of copying into Vec<u32> on init. The trie is now
  truly zero-copy from .rodata as documented. Add bounds check on
  base array index to prevent panics on corrupted trie data.

  Replace rayon with std::thread::scope for multi-file parallelism.
  Removes 5 transitive dependencies (crossbeam-deque, crossbeam-epoch,
  crossbeam-utils, either, rayon-core). Scoped threads are sufficient
  here — the parallel work is per-file tokenization, not fine-grained.

  Pre-compile --ignore glob patterns once instead of constructing a
  fancy_regex::Regex on every file path test. Avoids O(files × patterns)
  regex compilations during recursive directory scans.

Document the h|1 zero-sentinel invariant on all three FNV hash
functions (fnv_hash, fnv_hash_pair, fnv_hash_concat) in both build.rs
and src/frozen.rs.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
@eordano
eordano merged commit d6d4144 into main Feb 28, 2026
5 checks passed
@eordano
eordano deleted the rust-cli branch February 28, 2026 21:11
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant