Skip to content

fix: ignore bytecode when importing kernels - #847

Closed
drbh wants to merge 1 commit into
mainfrom
kernel-load-ignore-pyc-bytecode
Closed

drbh wants to merge 1 commit into
mainfrom
kernel-load-ignore-pyc-bytecode

Conversation

@drbh

@drbh drbh commented Sep 24, 2026

Copy link
Copy Markdown
Collaborator

this pr fixes kernels being able to run a .pyc in place of the signed source

tldr; hash_variant skips __pycache__ and *.pyc (the interpreter writes them after the first import so they cannot be in the digest), but the default loader will happily run an unchecked-hash .pyc (PEP 552) without looking at the .py. so the digest validates while code that was never signed is what actually runs

Reproduce:

first setup a venv

uv venv
uv pip install -e ./kernels "torch==2.12.*" numpy
source .venv/bin/activate

check the correctly signed relu kernel

kernels verify-signature kernels-community/relu 1

it outputs

✅ torch212-metal-aarch64-darwin: the metadata is correctly signed

now we can poison the cached __init__.py with an unchecked-hash .pyc, we can do this by getting the kernel with get_kernel, then compiling a print statement into pyc and placing it in the cache.

save this into a file

# `plant_pyc.py`

import py_compile, tempfile
from importlib.util import cache_from_source
from pathlib import Path
from kernels import get_kernel

init = Path(get_kernel("kernels-community/relu", version=1).__file__)
evil = Path(tempfile.mkdtemp()) / "__init__.py"
evil.write_text(init.read_text() + '\nprint("running unsigned pyc!!!")\n')
py_compile.compile(str(evil), cfile=cache_from_source(str(init)), dfile=str(init),
                   invalidation_mode=py_compile.PycInvalidationMode.UNCHECKED_HASH)
print("planted", cache_from_source(str(init)))

then run it

python plant_pyc.py

now we can recheck the signature of the relu kernel

kernels verify-signature kernels-community/relu 1

which outputs that it is still correctly signed

✅ torch212-metal-aarch64-darwin: the metadata is correctly signed

but if we run the kernel after planting the unchecked-hash .pyc, we see that the unsigned code executes

python -c 'from kernels import get_kernel; get_kernel("kernels-community/relu", version=1)'
running unsigned pyc!!!

Solution:

_import_from_path now loads the kernel with a source only loader, and registers the same loader for every directory in the kernel module so submodules are compiled from the .py too. sourceless .pyc files are not importable from the kernel at all. this only matters once verification is enforcing, but it keeps the digest covering everything that executes

note: since bytecode is never read, the kernel python is compiled on each load. this is paid once per process (the module is cached after that) and does not touch the .so loading or the triton cache in ~/.triton/cache. compiling every .py in the variant takes 0.13 ms for relu, 0.7 ms for triton-scaled-mm, 3.8 ms for triton-layer-norm, 8.6 ms for liger_kernels and 22 ms for triton_kernels (40 files, 242 KB), which is the worst case i found and small next to importing torch

@HuggingFaceDocBuilderDev

Copy link
Copy Markdown

The docs for this PR live here. All of your documentation changes will be reflected on that endpoint. The docs are available until 30 days after the last update.

@github-actions

Copy link
Copy Markdown

Coverage report — kernels/

Measured on: Python 3.10 / Torch 2.13.0.
Other CI configurations are not included in this number.
Hardware-gated code paths (ROCm/XPU/NPU/Darwin/Windows) are excluded or unreachable on the Linux+CUDA runner.

Total coverage: 87.2% — threshold: 80% — ✅

Per-file breakdown
Name Stmts Miss Cover Missing
src/kernels/__init__.py 14 0 100%
src/kernels/_system.py 6 1 83% 10
src/kernels/_versions.py 130 14 89% 53, 59-60, 63-64, 102, 165-170, 199, 219
src/kernels/archs.py 56 1 98% 94
src/kernels/backends.py 213 62 71% 42, 46, 50-53, 70, 92, 110, 119, 123, 127-129, 150, 159, 163, 167-169, 190, 201, 203, 210-213, 226, 230, 234-254, 262, 285-305
src/kernels/compat.py 9 1 89% 5
src/kernels/deps.py 70 1 99% 56
src/kernels/hf_hub.py 63 2 97% 21, 23
src/kernels/importer.py 58 6 90% 82, 115, 122, 125, 139-140
src/kernels/install.py 21 7 67% 76-100
src/kernels/layer/__init__.py 6 0 100%
src/kernels/layer/_interval_tree.py 103 4 96% 23, 52, 147, 150
src/kernels/layer/device.py 48 14 71% 42, 47-49, 91, 96-98, 101, 149, 152, 155-157
src/kernels/layer/func.py 85 6 93% 90, 115, 191, 311, 338, 368
src/kernels/layer/globals.py 5 0 100%
src/kernels/layer/kernelize.py 80 8 90% 258, 293, 301-302, 308, 312, 328-330
src/kernels/layer/layer.py 215 14 93% 182, 229, 256, 390, 470-471, 492, 500, 511, 540, 544, 557, 610, 640
src/kernels/layer/mode.py 14 0 100%
src/kernels/layer/repos.py 144 42 71% 27, 33, 36-43, 63-64, 70, 73-76, 90, 94, 103-104, 110, 113-116, 123-124, 130, 133-136, 143-144, 150, 153-156, 163-164, 170, 173-176, 257
src/kernels/load.py 71 2 97% 338, 378
src/kernels/locking.py 89 64 28% 35-83, 91-98, 102-125, 137, 152-159, 165-175, 179-186
src/kernels/python_deps.py 58 6 90% 59-60, 64-65, 101, 104
src/kernels/resolver.py 156 2 99% 220, 226
src/kernels/status.py 50 2 96% 25, 79
src/kernels/validate.py 88 5 94% 9, 100, 167, 190-191
src/kernels/variants.py 278 19 93% 65, 96, 117, 147, 256-257, 299-302, 304, 388-394, 400-406, 437-443, 455-461
src/kernels/verify.py 127 6 95% 46, 202-204, 318-319
TOTAL 2257 289 87%

Updated by the Test kernels workflow on commit 3efba49429dfb6a94abaea6c0ee2316fefa1ddd9.

return list(_loaded_kernels.values())


class _SourceOnlyLoader(importlib.machinery.SourceFileLoader):

@danieldk danieldk Sep 24, 2026 •

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

But we don't consider local cache poisoning as an attack vector. Once the cache could be poisoned, the attacker could also just replace .py files after all, or the program loading the kernel, or LD_PRELOAD a library to intercept calls, etc.

As long as we don't download .pyc files we should be good.

@drbh

drbh commented Sep 24, 2026

Copy link
Copy Markdown
Collaborator Author

closing as we verify the source but not artifacts that the user has generated since that extends to many cached items outside of the kernels libraries control. users are encouraged to use the verified source )(not cached files) where possible

@drbh drbh closed this Sep 24, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants