Skip to content

perf: add fast-path in token_next_by and eliminate list copy in token_index - #929

Open
av27-lgtm wants to merge 1 commit into
andialbrecht:masterfrom
av27-lgtm:perf/fastpath-token-matching
Open

av27-lgtm wants to merge 1 commit into
andialbrecht:masterfrom
av27-lgtm:perf/fastpath-token-matching

Conversation

@av27-lgtm

Copy link
Copy Markdown

Thanks for contributing!

Before submitting your pull request please have a look at the
following checklist:

  • ran the tests (pytest)
  • all style issues addressed (ruff)
  • your changes are covered by tests
  • your changes are documented, if needed

What this PR does

Two micro-optimizations in TokenList that together yield a measurable
throughput improvement on parse-heavy workloads.

1. token_next_by: fast-path for single-criterion lookups

token_next_by is one of the hottest functions in sqlparse — profiling
shows 69,800 calls for a single complex format(reindent=True) run.

The current implementation always creates a lambda closure and routes
through _token_matching → imt(), which checks all three branches
(i, m, t) on every token. However, ~95% of call sites pass exactly
one criterion
.

This PR adds an inline fast-path that, when only one of t, m, or i
is provided (and it's not a list), scans self.tokens directly — skipping
lambda creation, tuple wrapping, and multi-branch imt() dispatch.

When multiple criteria are passed (rare), the original _token_matching
fallback is used unchanged.

2. token_index: use list.index(token, start) instead of slice

The original start + self.tokens[start:].index(token) creates a
temporary list slice on every call. Passing start directly to
list.index() is semantically identical but avoids the copy.

Benchmark results

Python 3.14, 4 realistic SQL queries, 300 iterations × 2 ops (parse +
format), 5 runs, median:

Operation Before After Improvement
parse() only 111.4 ops/s 135.6 ops/s +21.7%
parse() + format(reindent=True) 288.7 calls/s 316.1 calls/s +9.5%

All 506 tests pass, ruff check clean.

Optimization by ARCHON EVO engine — autonomous performance analysis.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant