perf(normalizers): fast-path BertNormalizer for non-CJK input - #2383
Open
bsachart wants to merge 1 commit into
Open
perf(normalizers): fast-path BertNormalizer for non-CJK input#2383bsachart wants to merge 1 commit into
bsachart wants to merge 1 commit into
Conversation
Add an early return to do_handle_chinese_chars() that scans for any Chinese character before allocating. For the common English/Latin case this eliminates: - Vec<(char, isize)> allocation - for_each() closure dispatch - transform() + alignment reconstruction The only cost is a single preliminary chars().any() scan, which short-circuits on the first match for actual CJK input.
Author
Benchmark results
Single encode shows a statistically significant improvement. Batch encode is dominated by parallelism overhead so the normalizer cost is amortized away — expected. |
Collaborator
|
Ty! See #2119 as we are refactor the codebase! if you want to target this new branch! |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
do_handle_chinese_charsunconditionally allocates aVec<(char, isize)>, iterates every character, and callstransform()to rebuild alignments — even when the input contains no Chinese characters. This is wasted work for the common English/Latin case.Change
One early-return guard. No API changes, no new dependencies, no behavior change.
What it skips (for non-CJK input)
Vecallocationfor_each()closure dispatchtransform()+ alignment reconstructionTrade-off
CJK input pays for one extra
chars().any()scan that short-circuits on the first match. For inputs that are mostly CJK this is negligible relative to thetransform()cost that follows.