diff --git a/README.md b/README.md index 72e792a..e965cde 100644 --- a/README.md +++ b/README.md @@ -5,9 +5,9 @@ # Contact: qubitium@modelcloud.ai, x.com/qubitium --> -# PyPcre (Python PCRE2 Binding) +# PyPcre (Python PCRE2 Binding) ๐งฌ -Modern `nogil` Python bindings for the PCRE2 library with `stdlib.re` API compatibility. +Fast, free-threaded Python bindings for `PCRE2` with a stable `stdlib.re`-compatible API. โก
@@ -19,52 +19,92 @@ Modern `nogil` Python bindings for the PCRE2 library with `stdlib.re` API compat
-## Latest News
-* 03/22/2026 [0.2.15](https://github.com/ModelCloud/PyPcre/releases/tag/v0.2.15): Python 3.15 `re` compatibility (`prefixmatch`, `NOFLAG`)
-* 03/21/2026 [0.2.14](https://github.com/ModelCloud/PyPcre/releases/tag/v0.2.14): Python 3.14 compatibility
-* 03/02/2026 [0.2.11](https://github.com/ModelCloud/PyPcre/releases/tag/v0.2.11): Auto-detect `Visual Studio` in Windows environments during install and compile.
-* 02/24/2026 [0.2.10](https://github.com/ModelCloud/PyPcre/releases/tag/v0.2.10): Allow a `Visual Studio` (VS) compiler version check override via an environment variable.
-* 12/15/2025 [0.2.8](https://github.com/ModelCloud/PyPcre/releases/tag/v0.2.8): Fixed multi-arch Linux OS compatibility when both x86_64 and i386 `pcre2` libraries are installed.
-* 10/20/2025 [0.2.4](https://github.com/ModelCloud/PyPcre/releases/tag/v0.2.4): Removed the dependency on a system `python3-dev` package. `Python.h` will be downloaded optimistically from python.org when needed.
+## Latest News ๐
+* 04/13/2026 [0.3.0](https://github.com/ModelCloud/PyPcre/releases/tag/v0.3.0): Lower-overhead public `Match` objects, faster hot-path `match()` / `search()` / `fullmatch()` / `findall()`, and tighter free-threaded execution. โก
+* 03/22/2026 [0.2.15](https://github.com/ModelCloud/PyPcre/releases/tag/v0.2.15): Python 3.15 `re` compatibility (`prefixmatch`, `NOFLAG`) โ
+* 03/21/2026 [0.2.14](https://github.com/ModelCloud/PyPcre/releases/tag/v0.2.14): Python 3.14 compatibility ๐
+* 03/02/2026 [0.2.11](https://github.com/ModelCloud/PyPcre/releases/tag/v0.2.11): Auto-detect `Visual Studio` in Windows environments during install and compile. ๐ช
+* 02/24/2026 [0.2.10](https://github.com/ModelCloud/PyPcre/releases/tag/v0.2.10): Allow a `Visual Studio` (VS) compiler version check override via an environment variable. ๐งฐ
+* 12/15/2025 [0.2.8](https://github.com/ModelCloud/PyPcre/releases/tag/v0.2.8): Fixed multi-arch Linux OS compatibility when both x86_64 and i386 `pcre2` libraries are installed. ๐ง
+* 10/20/2025 [0.2.4](https://github.com/ModelCloud/PyPcre/releases/tag/v0.2.4): Removed the dependency on a system `python3-dev` package. `Python.h` will be downloaded optimistically from python.org when needed. ๐ฆ
* 10/12/2025 [0.2.3](https://github.com/ModelCloud/PyPcre/releases/tag/v0.2.3): ๐ค Full `GIL=0` compliance for Python >= 3.13T. Reduced cache thread contention. Improved performance across all APIs. Expanded CI test coverage. FreeBSD, Solaris, and Windows compatibility validated.
* 10/09/2025 [0.1.0](https://github.com/ModelCloud/PyPcre/releases/tag/v0.1.0): ๐ First release. Thread-safe, with auto JIT, auto pattern caching, and optimistic linking to the system library for fast installs.
-## Why PyPcre
+## Why PyPcre โก
-PyPcre is a modern PCRE2 binding designed to be both fast and thread-safe in a `GIL=0` world. In the era of the global interpreter lock, Python had real threads but often only limited concurrency, aside from a handful of low-level APIs and packages. As Python moves toward a fuller `GIL=0` design, true multi-threaded concurrency becomes practical and brings Python closer to parity with other modern languages.
+PyPcre pairs Python's familiar `re`-compatible API with the real `PCRE2` engine. You keep the ergonomics of the standard library while gaining a more capable regex engine, optional JIT, explicit threading support, and a binding designed and tested for free-threaded Python. ๐ง โก
-Many Python regular expression packages either segfault under `GIL=0` or suffer suboptimal performance because they were not designed with threaded execution in mind.
+### Big Wins ๐
-PyPcre is fully CI-tested. Every API and PCRE2 flag is exercised in a continuous development environment backed by the ModelCloud.AI team. Fuzz (clobber) tests are also run to catch memory safety, accuracy, and memory leak regressions.
+- ๐งฌ **Full power of PCRE2**: PyPcre uses the real `PCRE2` engine, so you get native compile options, semantics, JIT, and upstream tuning.
+- ๐ฅ **More expressive regex syntax**: `PCRE2` supports constructs beyond stdlib `re`, including atomic groups `(?>...)`, possessive quantifiers `++`, branch-reset groups `(?|...)`, richer lookarounds, and backtracking control verbs like `(*SKIP)(*FAIL)`.
+- ๐งต **Thread-safe into `nogil`**: PyPcre is built for `PYTHON_GIL=0`, with CI coverage, lock-aware caches, reusable match/JIT resources, and `parallel_map()` for multi-subject fan-out.
+- โก **Fast on real workloads**: `PCRE2` JIT plus cached compiled patterns lets PyPcre match or beat `re` and `regex` on many common scans, especially multiline searches, lookaround-heavy patterns, and free-threaded execution.
+- ๐ก๏ธ **Safer operational story**: PyPcre prefers the system `libpcre2-8` shared library so normal OS package updates can bring security and bug-fix benefits without a bundled fork.
+- โ
**Validated thoroughly**: the project runs API tests, fuzz tests, memory-safety checks, local `valgrind` leak checks, and `massif` heap profiles. Recent local profiling found `0` definite leaks and `0` possible leaks in both the public API and raw binding paths.
-For safety, PyPcre preferentially links against the OS-provided `libpcre2` package so it can benefit from upstream security patches. You can force a full source build with the `PYPCRE_BUILD_FROM_SOURCE=1` environment variable.
+### Quick Comparison ๐ฅ
-## Installation
+| Area | PyPcre | `stdlib.re` | `regex` |
+| --- | --- | --- | --- |
+| Engine | Full `PCRE2` โ
| CPython stdlib engine | Separate engine, not `PCRE2` |
+| `PCRE2` syntax and flags | Full access โ
| No | No |
+| Syntax power | Very rich โ
| More limited | Rich, but different from `PCRE2` |
+| JIT execution | `PCRE2` JIT โ
| No | No |
+| `re`-compatible API surface | Stable and familiar โ
| Native | Similar, but not the main goal |
+| Free-threaded support | Built and tested for `PYTHON_GIL=0` โ
| No explicit PyPcre-style layer | Not a project focus here |
+| Built-in threaded subject fan-out | `parallel_map()` โ
| No | No |
+| System library updates | Uses system `libpcre2-8` by default โ
| N/A | N/A |
+
+### Benchmark Highlights ๐
+
+Measured on a `Python 3.14.3` free-threaded build on x86_64 Linux with compiled-pattern reuse. Times are best-of-5; lower is better.
+
+| Workload | Operation | PyPcre | `re` | `regex` | PyPcre edge |
+| --- | --- | ---: | ---: | ---: | --- |
+| First `ERROR` line in a multiline log buffer | `search` | `3.68 ms` | `51.72 ms` | `5.67 ms` | `14.0x` vs `re`, `1.54x` vs `regex` |
+| Extract only `WARN` / `ERROR` lines | `findall` | `6.41 ms` | `91.84 ms` | `91.14 ms` | `14.3x` vs `re`, `14.2x` vs `regex` |
+| Per-line full-name extraction | `findall` | `22.28 ms` | `172.38 ms` | `218.29 ms` | `7.74x` vs `re`, `9.80x` vs `regex` |
+| Lookbehind + negative-lookahead extraction | `findall` | `50.23 ms` | `53.35 ms` | `57.03 ms` | `1.06x` vs `re`, `1.14x` vs `regex` |
+| UUID extraction | `findall` | `77.49 ms` | `83.19 ms` | `134.87 ms` | `1.07x` vs `re`, `1.74x` vs `regex` |
+| Boundary-aware token scan | `findall` | `127.76 ms` | `128.03 ms` | `146.37 ms` | effectively tied with `re`, `1.15x` vs `regex` |
+
+### Free-Threaded Benchmark Highlights ๐งต
+
+Measured in the same environment with `8` threads sharing one compiled pattern. Times are best-of-3; lower is better.
+
+| Workload | Threads | PyPcre | `re` | `regex` | PyPcre edge |
+| --- | ---: | ---: | ---: | ---: | --- |
+| First `ERROR` line in a multiline log buffer | `8` | `25.34 ms` | `38.83 ms` | `40.34 ms` | `1.53x` vs `re`, `1.59x` vs `regex` |
+| Extract only `WARN` / `ERROR` lines | `8` | `28.58 ms` | `65.54 ms` | `73.55 ms` | `2.29x` vs `re`, `2.57x` vs `regex` |
+| Per-line full-name extraction | `8` | `31.68 ms` | `123.44 ms` | `164.80 ms` | `3.90x` vs `re`, `5.20x` vs `regex` |
+
+PyPcre is the stronger all-around choice when you want more than the baseline: full `PCRE2` features, more expressive syntax, JIT, explicit free-threaded support, and a stable `re`-compatible API surface. It keeps Python ergonomics while giving you a substantially more capable engine. ๐
+
+## Installation ๐ฆ
```bash
pip install PyPcre
```
-The package prefers linking against the system `libpcre2-8` shared library for fast installs and to inherit security updates from the OS. See [Building](#building) for manual build details.
+By default, the package links against the system `libpcre2-8` shared library for fast installs and to inherit OS security updates. See [Building](#building) for manual build details.
-## Platform Support (Validated)
+## Platform Support (Validated) โ
`Linux`, `macOS`, `Windows`, `WSL`, `FreeBSD`
-## Usage
+## Usage ๐ ๏ธ
-If you already rely on the standard library `re`, migrating is as
-simple as changing your import:
+If you already use the standard library `re`, migration is often just an import swap:
```python
import pcre as re
```
-The high-level API keeps the standard library shape, so most existing `re`
-code can move over with little or no rewriting.
+The high-level API stays close to the standard library, so most existing `re` code can move over with little or no rewriting.
-### Quick start
+### Quick start ๐
```python
from pcre import compile, findall, match, search, Flag
@@ -76,7 +116,7 @@ pattern = compile(rb"\d+", flags=Flag.MULTILINE)
numbers = pattern.findall(b"line 1\nline 22")
```
-### User-facing API
+### API Overview ๐งญ
- Module helpers: `prefixmatch`, `match`, `search`, `fullmatch`, `finditer`,
`findall`, `split`, `sub`, `subn`, `compile`, `escape`, `purge`, and
@@ -94,7 +134,7 @@ numbers = pattern.findall(b"line 1\nline 22")
- Errors are raised as `pcre.PcreError`; `error` and `PatternError` are kept as
compatibility aliases.
-### Common examples
+### Common examples ๐งช
Compiled patterns:
@@ -124,7 +164,7 @@ pattern = compile(br"\w+")
print(pattern.findall(b"ab cd")) # [b'ab', b'cd']
```
-### Stdlib `re` compatibility
+### Stdlib `re` compatibility ๐
- Module-level helpers and the `Pattern` class follow the same call shapes as
the standard library `re` module, including `pos`, `endpos`, and `flags`
@@ -148,7 +188,7 @@ print(pattern.findall(b"ab cd")) # [b'ab', b'cd']
patterns so escaping semantics remain identical.
- String patterns enable Unicode behavior by default. Byte patterns do not.
-### `regex` package compatibility
+### `regex` package compatibility ๐
The [`regex`](https://pypi.org/project/regex/) package interprets
`\uXXXX` and `\UXXXXXXXX` escapes as UTF-8 code points, while PCRE2 expects
@@ -166,7 +206,7 @@ Set the default behavior globally with `pcre.configure(compat_regex=True)`
so that subsequent calls to `compile()` and the module-level helpers apply
the conversion without repeating the flag.
-### Common issues
+### Common issues โ ๏ธ
- Unsupported stdlib flags such as `re.DEBUG`, `re.LOCALE`, and `re.ASCII`
raise `ValueError`. If you want ASCII-style behavior, use `pcre.ASCII` or
@@ -178,7 +218,7 @@ the conversion without repeating the flag.
- Most users do not need to tune caching, JIT, or threading. The defaults are
intended to work well out of the box.
-### Optional runtime controls
+### Optional runtime controls ๐๏ธ
- `pcre.configure(jit=False)` disables JIT globally. `Flag.JIT` and
`Flag.NO_JIT` let you override that per pattern.
@@ -188,10 +228,9 @@ the conversion without repeating the flag.
`shutdown_thread_pool()`, `Flag.THREADS`, and `Flag.NO_THREADS` are available
if you want to opt into or restrict threaded execution.
-## Building
+## Building ๐๏ธ
-The extension links against an existing PCRE2 installation (the `libpcre2-8`
-variant). Install the development headers for your platform before building,
+The extension links against an existing `libpcre2-8` installation. Install the development headers for your platform before building,
for example `apt install libpcre2-dev` on Debian/Ubuntu, `dnf install pcre2-devel`
on Fedora/RHEL derivatives, or `brew install pcre2` on macOS.
diff --git a/pcre/pcre.py b/pcre/pcre.py
index 3a98efe..84b4880 100644
--- a/pcre/pcre.py
+++ b/pcre/pcre.py
@@ -28,7 +28,7 @@
NO_UTF: int = int(Flag.NO_UTF)
NO_UCP: int = int(Flag.NO_UCP)
from .re_compat import (
- Match,
+ Match as _CompatMatch,
TemplatePatternStub,
coerce_group_value,
coerce_subject_slice,
@@ -53,6 +53,9 @@
_CPattern = _pcre2.Pattern
PcreError = _pcre2.PcreError
+Match = getattr(_pcre2, "Match", _CompatMatch)
+_ATTACH_MATCH = getattr(_pcre2, "_attach_match", None)
+_RAW_MATCH_TYPE = getattr(_pcre2, "Match", None)
FlagInput = int | _std_re.RegexFlag | Iterable[int | _std_re.RegexFlag]
@@ -65,6 +68,14 @@
_THREAD_MODE_AUTO = "auto"
+def _can_attach_match(raw: Any) -> bool:
+ return (
+ _ATTACH_MATCH is not None
+ and _RAW_MATCH_TYPE is not None
+ and isinstance(raw, _RAW_MATCH_TYPE)
+ )
+
+
def _resolve_jit_setting(jit: bool | None) -> bool:
if jit is None:
return _DEFAULT_JIT
@@ -162,13 +173,6 @@ def _normalise_flags(flags: FlagInput) -> int:
raise TypeError("flags must be an int, stdlib re flag, or an iterable of those")
-def _call_with_optional_end(method, subject: Any, pos: int, endpos: int | None, options: int):
- resolved_end = resolve_endpos(subject, endpos)
- if endpos is None:
- return method(subject, pos=pos, options=options), resolved_end
- return method(subject, pos=pos, endpos=resolved_end, options=options), resolved_end
-
-
class Pattern:
"""High-level wrapper around the C-backed :class:`pcre_ext_c.Pattern`."""
@@ -223,8 +227,10 @@ def enable_auto_threads(self) -> None:
self._thread_mode = _THREAD_MODE_AUTO
def _update_group_hint(self, match: Match) -> None:
+ if self._groups_hint is not None:
+ return
groups_count = len(match.groups())
- if self._groups_hint is None or groups_count > self._groups_hint:
+ if groups_count > 0:
self._groups_hint = groups_count
def _wrap_match(
@@ -236,7 +242,9 @@ def _wrap_match(
) -> Match | None:
if raw is None:
return None
- wrapped = Match(self, raw, subject, pos, end_boundary)
+ if _can_attach_match(raw):
+ return _ATTACH_MATCH(raw, self)
+ wrapped = _CompatMatch(self, raw, subject, pos, end_boundary)
self._update_group_hint(wrapped)
return wrapped
@@ -248,8 +256,22 @@ def match(
endpos: int | None = None,
options: int = 0,
) -> Match | None:
- subject = prepare_subject(subject)
- raw, resolved_end = _call_with_optional_end(self._pattern.match, subject, pos, endpos, options)
+ if type(subject) is memoryview:
+ subject = subject.tobytes()
+ if endpos is None:
+ raw = self._pattern.match(subject, pos=pos, options=options)
+ if raw is None:
+ return None
+ if _can_attach_match(raw):
+ return _ATTACH_MATCH(raw, self)
+ resolved_end = len(subject)
+ else:
+ resolved_end = resolve_endpos(subject, endpos)
+ raw = self._pattern.match(subject, pos=pos, endpos=resolved_end, options=options)
+ if raw is None:
+ return None
+ if _can_attach_match(raw):
+ return _ATTACH_MATCH(raw, self)
return self._wrap_match(raw, subject, pos, resolved_end)
prefixmatch = match
@@ -262,8 +284,22 @@ def search(
endpos: int | None = None,
options: int = 0,
) -> Match | None:
- subject = prepare_subject(subject)
- raw, resolved_end = _call_with_optional_end(self._pattern.search, subject, pos, endpos, options)
+ if type(subject) is memoryview:
+ subject = subject.tobytes()
+ if endpos is None:
+ raw = self._pattern.search(subject, pos=pos, options=options)
+ if raw is None:
+ return None
+ if _can_attach_match(raw):
+ return _ATTACH_MATCH(raw, self)
+ resolved_end = len(subject)
+ else:
+ resolved_end = resolve_endpos(subject, endpos)
+ raw = self._pattern.search(subject, pos=pos, endpos=resolved_end, options=options)
+ if raw is None:
+ return None
+ if _can_attach_match(raw):
+ return _ATTACH_MATCH(raw, self)
return self._wrap_match(raw, subject, pos, resolved_end)
def fullmatch(
@@ -274,8 +310,22 @@ def fullmatch(
endpos: int | None = None,
options: int = 0,
) -> Match | None:
- subject = prepare_subject(subject)
- raw, resolved_end = _call_with_optional_end(self._pattern.fullmatch, subject, pos, endpos, options)
+ if type(subject) is memoryview:
+ subject = subject.tobytes()
+ if endpos is None:
+ raw = self._pattern.fullmatch(subject, pos=pos, options=options)
+ if raw is None:
+ return None
+ if _can_attach_match(raw):
+ return _ATTACH_MATCH(raw, self)
+ resolved_end = len(subject)
+ else:
+ resolved_end = resolve_endpos(subject, endpos)
+ raw = self._pattern.fullmatch(subject, pos=pos, endpos=resolved_end, options=options)
+ if raw is None:
+ return None
+ if _can_attach_match(raw):
+ return _ATTACH_MATCH(raw, self)
return self._wrap_match(raw, subject, pos, resolved_end)
def finditer(
@@ -286,7 +336,8 @@ def finditer(
endpos: int | None = None,
options: int = 0,
) -> Generator[Match, None, None]:
- subject = prepare_subject(subject)
+ if type(subject) is memoryview:
+ subject = subject.tobytes()
origin_pos = pos
resolved_end = resolve_endpos(subject, endpos)
backend_iter = getattr(self._pattern, "finditer", None)
@@ -298,9 +349,7 @@ def finditer(
raw_iter = None
if raw_iter is not None:
for raw in raw_iter:
- match_obj = Match(self, raw, subject, origin_pos, resolved_end)
- self._update_group_hint(match_obj)
- yield match_obj
+ yield self._wrap_match(raw, subject, origin_pos, resolved_end)
return
search_end = resolved_end if endpos is not None else -1
@@ -312,8 +361,7 @@ def finditer(
if raw is None:
break
- match_obj = Match(self, raw, subject, origin_pos, resolved_end)
- self._update_group_hint(match_obj)
+ match_obj = self._wrap_match(raw, subject, origin_pos, resolved_end)
yield match_obj
start, end = match_obj.span()
@@ -334,6 +382,25 @@ def findall(
endpos: int | None = None,
options: int = 0,
) -> List[Any]:
+ if type(subject) is memoryview:
+ subject = subject.tobytes()
+ backend_iter = getattr(self._pattern, "finditer", None)
+ if backend_iter is not None:
+ compiled_end = -1 if endpos is None else resolve_endpos(subject, endpos)
+ try:
+ raw_iter = backend_iter(subject, pos=pos, endpos=compiled_end, options=options)
+ except TypeError:
+ raw_iter = None
+ if raw_iter is not None:
+ results: List[Any] = []
+ for raw in raw_iter:
+ groups = raw.groups()
+ if groups:
+ results.append(groups[0] if len(groups) == 1 else groups)
+ else:
+ results.append(raw.group(0))
+ return results
+
results: List[Any] = []
for match_obj in self.finditer(subject, pos=pos, endpos=endpos, options=options):
groups = match_obj.groups()
diff --git a/pcre/re_compat.py b/pcre/re_compat.py
index c36653b..2a7ece9 100644
--- a/pcre/re_compat.py
+++ b/pcre/re_compat.py
@@ -151,6 +151,24 @@ def render_template(parsed: Any, match: "Match", *, is_bytes: bool, empty: Any)
return join_parts(pieces, is_bytes=is_bytes)
+def expand_match_template(match: Any, template: Any) -> Any:
+ is_bytes = is_bytes_like(match.string)
+ empty = b"" if is_bytes else ""
+ if is_bytes:
+ if not is_bytes_like(template):
+ raise TypeError("template must be bytes-like for bytes matches")
+ template = bytes(template)
+ else:
+ if not isinstance(template, str):
+ raise TypeError("template must be str for text matches")
+
+ parsed = _parser.parse_template(
+ template,
+ TemplatePatternStub(match.re.groups, match.re.groupindex),
+ )
+ return render_template(parsed, match, is_bytes=is_bytes, empty=empty)
+
+
def maybe_infer_group_count(pattern_source: Any) -> int | None:
normalised = pattern_source
if isinstance(normalised, memoryview):
@@ -270,21 +288,7 @@ def span(self, group: Any = 0) -> tuple[int, int]:
return self._match.span(group)
def expand(self, template: Any) -> Any:
- is_bytes = is_bytes_like(self._string)
- empty = b"" if is_bytes else ""
- if is_bytes:
- if not is_bytes_like(template):
- raise TypeError("template must be bytes-like for bytes matches")
- template = bytes(template)
- else:
- if not isinstance(template, str):
- raise TypeError("template must be str for text matches")
-
- parsed = _parser.parse_template(
- template,
- TemplatePatternStub(self.re.groups, self.re.groupindex),
- )
- return render_template(parsed, self, is_bytes=is_bytes, empty=empty)
+ return expand_match_template(self, template)
@property
def re(self) -> Any:
diff --git a/pcre_ext/pcre2.c b/pcre_ext/pcre2.c
index bd89b76..a93aa84 100644
--- a/pcre_ext/pcre2.c
+++ b/pcre_ext/pcre2.c
@@ -27,9 +27,11 @@ resolve_pcre2_prerelease(void)
return raw;
}
+/* Process-wide library metadata cached once during module initialization. */
static char pcre2_library_version[64] = "unknown";
static ATOMIC_VAR(int) pcre2_version_initialized = 0;
#if defined(PCRE2_USE_OFFSET_LIMIT)
+/* -1 unknown, 0 unsupported, 1 supported by the loaded PCRE2 runtime. */
static ATOMIC_VAR(int) offset_limit_support = ATOMIC_VAR_INIT(-1);
#endif
@@ -162,6 +164,7 @@ static void
Match_dealloc(MatchObject *self)
{
Py_XDECREF(self->pattern);
+ Py_XDECREF(self->public_pattern);
Py_XDECREF(self->subject);
Py_XDECREF(self->utf8_owner);
pcre_free(self->ovector);
@@ -176,6 +179,64 @@ Match_repr(MatchObject *self)
return PyUnicode_FromFormat("