Skip to content

fix(trees): improve tree node search performance - #8564

Draft
grantfitzsimmons wants to merge 14 commits into
mainfrom
issue-7751
Draft

grantfitzsimmons wants to merge 14 commits into
mainfrom
issue-7751

Conversation

@grantfitzsimmons

@grantfitzsimmons grantfitzsimmons commented Sep 21, 2026 •

Copy link
Copy Markdown
Member

Begins to fix #7751, but much to do

This makes it so much faster to query based on a particular rank. For example, querying Collection Object using "Heterelmis" in the KU Entomology taxon tree would take many minutes to complete (just fetching the first 40), whereas with these changes it is nearly instantaneous (see /specify/query/fromtree/taxon/2004/ when logged into the database in the KUEntoPinned collection).

Important

These changes were largely written using OpenAI’s Codex with GPT6-Astra. Codex inspected the Query Builder implementation and generated SQL, identified bottlenecks in tree queries, implemented the optimization, wrote regression tests, and ran local benchmarks and the test suite. This work tests these tools’ potential to contribute code for integration into Specify.

Description

The Query Builder tree rank filters currently resolve ancestors by joining the tree table to itself once for every possible rank. A taxonomy with 22 ranks could therefore add 21 parent joins (like KU ento tree), even when the query only filtered on one rank. This made otherwise selective Taxon, Geography, and Collection Object queries slow and affected pages, counts, and exports that share the query compiler.

This benefit will mostly be felt when querying directly on a tree name (which is a very common workflow). See below:

Phellopsis.mov

GPT6-Astra:
The following measurements used repeated executions of 20-row Taxon queries on
the local development dataset:

Genus filter Before: parent walk After Observed change
Exact (Phellopsis) 21 joins, 1.35–1.45 s 1 interval join, 34–49 ms 28–43× faster
Contains (ell) 21 joins, 101–105 ms 3 recursive-lookup joins, 77–92 ms 1.1–1.4× faster
Starts With (P%) 21 joins, 67–68 ms 3 recursive-lookup joins, 71–79 ms 4–18% slower

This change selects the tree SQL strategy from the filter semantics:

  • Exact rank matches join the matching ancestor once and use its maintained NodeNumber/HighestChildNodeNumber interval to find descendants.
  • Pattern and other positive scalar matches build one recursive descendant-to-rank lookup and join the resolved ancestor by primary key.
  • Empty, negated, relationship, and date filters retain the existing parent walk so their null and missing-ancestor behavior is unchanged. This should likely be addressed in a future update.

Tree definitions and rank items are loaded in bulk and reused while compiling a query. Missing or reversed node-number intervals automatically fall back to the existing parent-walk behavior. Saved-query data and API contracts are unchanged.

On the development dataset, an exact Genus filter dropped from 21 parent joins and 1.35–1.45 seconds to one interval join and 34–49 ms. Contains and prefix filters used three joins and completed in 70–92 ms. Results are compared against the previous parent-walk implementation in the regression suite.

There is no migration, new index, dependency change, worker change, localization change, or user-visible configuration change.

Checklist

  • Self-review the PR after opening it to make sure the changes look good and self-explanatory (or properly documented)
  • Add relevant issue to release milestone
  • Add PR to documentation list
  • Add automated tests
  • Add a reverse migration if a migration is present in the PR (not applicable; this PR has no migration)
  • Add migration function to run_key_migration_functions.py (not applicable; this PR has no migration)

Testing instructions

This is easiest to test on databases with large trees (e.g., Taxon in herb-rbge or KU entomology), but we need to test on existing databases.

  • Go to a tree node with over 1,000 records pointing to it (e.g., Heterelmis in KU Ento)
  • Build queries that filter Taxon and Geography ranks using exact, prefix, and contains operators. Check the first page, count, and export results against the same queries on the latest release of v7.
  • Verify an empty or negated rank filter and a query restricted by a record set. Results should be unchanged while positive rank filters should avoid the full chain of parent self-joins.

Summary by CodeRabbit

  • Performance

    • Improved search performance for supported tree-based filters, especially when querying large hierarchies.
    • Reduced repeated metadata lookups during tree filtering.
  • Bug Fixes

    • Improved handling of tree filters across IDs, ranges, negation, geography, boolean, and synonym-based searches.
    • Added reliable fallback behavior when hierarchy numbering is unavailable or inconsistent.
  • Tests

    • Expanded coverage to verify optimized tree searches return the same results as existing filtering behavior.

Use live node-number ranges for rank ID equality and reusable recursive rank lookups for other positive scalar filters. Preserve parent lookups for negative/empty filters, incomplete numbering, recordsets, and explicit record ID selections. Cover scoping, preferred taxa, OR combinations, tree moves, and result equivalence.
Use live node-number ranges for rank ID equality and reusable recursive rank lookups for other positive scalar filters. Preserve parent lookups for negative/empty filters, incomplete numbering, recordsets, and explicit record ID selections. Cover scoping, preferred taxa, OR combinations, tree moves, and result equivalence.
@grantfitzsimmons grantfitzsimmons added the 4 - Performance Issues related to performance, concurrency, and optimization label Sep 21, 2026
@github-project-automation github-project-automation Bot moved this to 📋Back Log in General Tester Board Sep 21, 2026
@coderabbitai

coderabbitai Bot commented Sep 21, 2026 •

Copy link
Copy Markdown
Contributor

Review Change StackReview Change Stack

Important

Draft PR not reviewed

Draft PRs are not automatically reviewed by default.

  • Trigger a manual review

To automatically review draft PRs, update your CodeRabbit configuration:

reviews:
  auto_review:
    drafts: true
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Commit to this branch
  • Create a new PR

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@grantfitzsimmons grantfitzsimmons added 2 - Trees Issues that are related to the tree system and related functionalities. 2 - Queries Issues that are related to the query builder or queries in general labels Sep 21, 2026
Comment thread specifyweb/backend/stored_queries/queryfield.py Fixed

logger = logging.getLogger(__name__)
if TYPE_CHECKING:
from .queryfieldspec import QueryFieldSpec

@classmethod
def from_spqueryfield(cls, field: EphemeralField, value: str | None=None):
from .queryfieldspec import QueryFieldSpec

def add_to_query(self, query, no_filter=False, formatauditobjs=False, collection=None, user=None):
def add_to_query(self, query, no_filter=False, formatauditobjs=False, collection=None, user=None, optimize_tree=True):
from .queryfieldspec import TreeRankQuery
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

2 - Queries Issues that are related to the query builder or queries in general 2 - Trees Issues that are related to the tree system and related functionalities. 4 - Performance Issues related to performance, concurrency, and optimization

Projects

Status: 📋Back Log

Development

Successfully merging this pull request may close these issues.

[Large Databases]: Improve large tree Query Builder performance

2 participants