fix(trees): improve tree node search performance - #8564
Draft
grantfitzsimmons wants to merge 14 commits into
Draft
grantfitzsimmons wants to merge 14 commits into
grantfitzsimmons wants to merge 14 commits into
Conversation
Use live node-number ranges for rank ID equality and reusable recursive rank lookups for other positive scalar filters. Preserve parent lookups for negative/empty filters, incomplete numbering, recordsets, and explicit record ID selections. Cover scoping, preferred taxa, OR combinations, tree moves, and result equivalence.
Use live node-number ranges for rank ID equality and reusable recursive rank lookups for other positive scalar filters. Preserve parent lookups for negative/empty filters, incomplete numbering, recordsets, and explicit record ID selections. Cover scoping, preferred taxa, OR combinations, tree moves, and result equivalence.
Contributor
|
Important Draft PR not reviewedDraft PRs are not automatically reviewed by default.
To automatically review draft PRs, update your CodeRabbit configuration: reviews:
auto_review:
drafts: true✨ Finishing Touches🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
|
|
||
| logger = logging.getLogger(__name__) | ||
| if TYPE_CHECKING: | ||
| from .queryfieldspec import QueryFieldSpec |
|
|
||
| @classmethod | ||
| def from_spqueryfield(cls, field: EphemeralField, value: str | None=None): | ||
| from .queryfieldspec import QueryFieldSpec |
|
|
||
| def add_to_query(self, query, no_filter=False, formatauditobjs=False, collection=None, user=None): | ||
| def add_to_query(self, query, no_filter=False, formatauditobjs=False, collection=None, user=None, optimize_tree=True): | ||
| from .queryfieldspec import TreeRankQuery |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Begins to fix #7751, but much to do
This makes it so much faster to query based on a particular rank. For example, querying
Collection Object using "Heterelmis"in the KU Entomology taxon tree would take many minutes to complete (just fetching the first 40), whereas with these changes it is nearly instantaneous (see/specify/query/fromtree/taxon/2004/when logged into the database in theKUEntoPinnedcollection).Important
These changes were largely written using OpenAI’s Codex with GPT6-Astra. Codex inspected the Query Builder implementation and generated SQL, identified bottlenecks in tree queries, implemented the optimization, wrote regression tests, and ran local benchmarks and the test suite. This work tests these tools’ potential to contribute code for integration into Specify.
Description
The Query Builder tree rank filters currently resolve ancestors by joining the tree table to itself once for every possible rank. A taxonomy with 22 ranks could therefore add 21 parent joins (like KU ento tree), even when the query only filtered on one rank. This made otherwise selective Taxon, Geography, and Collection Object queries slow and affected pages, counts, and exports that share the query compiler.
This benefit will mostly be felt when querying directly on a tree name (which is a very common workflow). See below:
Phellopsis.mov
This change selects the tree SQL strategy from the filter semantics:
NodeNumber/HighestChildNodeNumberinterval to find descendants.Tree definitions and rank items are loaded in bulk and reused while compiling a query. Missing or reversed node-number intervals automatically fall back to the existing parent-walk behavior. Saved-query data and API contracts are unchanged.
On the development dataset, an exact Genus filter dropped from 21 parent joins and 1.35–1.45 seconds to one interval join and 34–49 ms. Contains and prefix filters used three joins and completed in 70–92 ms. Results are compared against the previous parent-walk implementation in the regression suite.
There is no migration, new index, dependency change, worker change, localization change, or user-visible configuration change.
Checklist
run_key_migration_functions.py(not applicable; this PR has no migration)Testing instructions
This is easiest to test on databases with large trees (e.g., Taxon in
herb-rbgeor KU entomology), but we need to test on existing databases.Heterelmisin KU Ento)v7.Summary by CodeRabbit
Performance
Bug Fixes
Tests