Conversation
Remove section comments, deduplicate CSS rules, compress DOM refs with a $ helper, consolidate event handlers, simplify diff rendering, and clean up build scripts. Bump version to 0.99.0 and add --version flag to Node.js CLI. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
|
Three tokenizer backends implemented from scratch (no tokenizer libraries): double-array trie for Claude, FNV-1a frozen hash tables for OpenAI tiktoken (o200k_base), and BPE with frozen hash tables for 7 HuggingFace models. All model data is embedded at build time via include_bytes! for zero-copy runtime access. Includes build.rs for compile-time data generation, flake.nix integration with rustModelsDir, and updated documentation. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
Tiktoken's byte-level BPE was O(n²): linear scan for the best merge pair + Vec::remove on every iteration. An adversarial input like "abcdefghijklmnopqrstuvwxyz"×20K (matched as a single regex piece) took 85s at 500K bytes, scaling quadratically. Replace with the same priority-queue + linked-list skip structure used by the HF BPE path: BinaryHeap<(rank, index, generation)> for O(log n) pops, next/prev/alive arrays for O(1) neighbor traversal, generation counters for cheap stale-entry invalidation with rank re-verification as a correctness backstop. Expected ~31x speedup at 500K based on the HF BPE path benchmarks under the same adversarial load. Add frozen_map_get_concat() and fnv_hash_concat() to frozen.rs — zero-allocation lookup of two concatenated slices in a frozen hash table, avoiding the temporary Vec that the old bpe_count needed for every pair probe. Remove the now-unused frozen_map_get(). Three other wins stacked in: Frozen hash tables: drop the power-of-2 slot count requirement. Replace bitmask indexing with Lemire fast range reduction (128-bit multiply + shift, no division). Slot counts shrink from next_power_of_two(⌈4n/3⌉) to exactly ⌈4n/3⌉ — ~35% fewer slots for large models, directly reducing embedded binary size since each slot is 14–18 bytes × ~200K entries. Claude double-array trie: reference TRIE_BIN as &'static [u8] slices instead of copying into Vec<u32> on init. The trie is now truly zero-copy from .rodata as documented. Add bounds check on base array index to prevent panics on corrupted trie data. Replace rayon with std::thread::scope for multi-file parallelism. Removes 5 transitive dependencies (crossbeam-deque, crossbeam-epoch, crossbeam-utils, either, rayon-core). Scoped threads are sufficient here — the parallel work is per-file tokenization, not fine-grained. Pre-compile --ignore glob patterns once instead of constructing a fancy_regex::Regex on every file path test. Avoids O(files × patterns) regex compilations during recursive directory scans. Document the h|1 zero-sentinel invariant on all three FNV hash functions (fnv_hash, fnv_hash_pair, fnv_hash_concat) in both build.rs and src/frozen.rs. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Three tokenizer backends implemented from scratch (no tokenizer
libraries): double-array trie for Claude, FNV-1a frozen hash tables
for OpenAI tiktoken (o200k_base), and BPE with frozen hash tables
for 7 HuggingFace models. All model data is embedded at build time
via include_bytes! for zero-copy runtime access.