Look inside an AI model file, and check it for hidden code
Drop a model downloaded from a hub to see what it is before any program loads it: a GGUF's architecture, context length, quantisation, tokenizer and chat template; every tensor's name, shape and type in GGUF, safetensors and PyTorch files; and, for the formats that are Python pickles (.pt, .pth, .bin, .ckpt, .pkl, .joblib), every import loading the file would perform, with anything like os.system flagged in red.
The file is read on your device a few megabytes at a time, so a 40 GB model is fine. Nothing is uploaded, and nothing in it is loaded or run.
What it shows
- GGUF (.gguf), the llama.cpp, Ollama and LM Studio format: the GGUF version, every metadata key and value (architecture, context length, embedding size, layers and heads, the quantisation from
general.file_typesuch as Q4_K_M, the tokenizer model and vocabulary size), the chat template in full as text, and a table of every tensor with its shape, ggml type (F16, Q8_0, Q4_K, IQ2_XS…) and size, plus the total parameter count. Long arrays such as the 32,000-entry vocabulary are shown as their count and first items. - safetensors (.safetensors): the
__metadata__block, every tensor with its dtype, shape, byte count and offsets, the total parameters, and a check that the offsets tile the data exactly — starting at 0, no gaps, no overlaps, ending at the end of the file. - PyTorch (.pt, .pth, .bin, .ckpt), both the ZIP that
torch.savehas written since PyTorch 1.6 and the older five-pickle stream, and raw pickles (.pkl, .pickle, uncompressed .joblib): everyGLOBAL,STACK_GLOBALandINSTimport in the pickle, whether it is called, the number of REDUCE (call), NEWOBJ (build an object) and BUILD (set its state) opcodes, and a tensor table read from the storage records PyTorch writes. - NumPy (.npy, .npz): dtype, shape and memory order of each array. An array of dtype
objectis really a pickle, and its imports are walked too.
Why a model file can run code
A .pt, .pth, .ckpt or older pytorch_model.bin is a Python pickle, and so are .pkl and most .joblib files. A pickle is not data: it is a small program for a stack machine that Python's pickle.load runs to rebuild objects. Its GLOBAL and STACK_GLOBAL instructions import any function by name, and REDUCE calls it with arguments from the file. A normal checkpoint uses that to call torch._utils._rebuild_tensor_v2 and collections.OrderedDict. A booby-trapped one calls os.system instead, the moment you write torch.load("model.pt").
Here is the whole of such a file, 95 bytes, as this page's own test writes it (the command is a harmless echo; the 94 MEMOIZE byte after each step is left out):
80 04 PROTO 4
95 54 00 … FRAME of 84 bytes
8c 02 "os" SHORT_BINUNICODE 'os'
8c 06 "system" SHORT_BINUNICODE 'system'
93 STACK_GLOBAL → imports os.system
8c 3c "echo …" SHORT_BINUNICODE the shell command
85 TUPLE1 → ('echo …',)
52 REDUCE → calls os.system('echo …')
2e STOP
This page reads those opcodes the way pickletools.dis does, with a hand-written walker that only records what each one would do, and reports: This file calls os.system when loaded: it runs a shell command on your computer. (Python on Linux and macOS pickles os.system as posix.system, and on Windows as nt.system; all three are flagged.) Imports are checked against a safe list close to the one PyTorch's own weights_only=True loader allows: the tensor-rebuilding helpers in torch._utils, torch storage classes and dtypes, collections.OrderedDict, NumPy's array reconstruction, and harmless builtins such as set and frozenset. Anything else is shown in red with a sentence saying what it does: shell commands, starting programs, evaluating code, touching files, opening network connections, or simply loading a class (a whole pickled model rather than its weights).
Since PyTorch 2.6 (January 2025), torch.load defaults to weights_only=True and refuses imports outside its list, and NumPy has refused object arrays in np.load by default since 1.16.3. Older versions, code that passes weights_only=False or allow_pickle=True, pickle.load and joblib.load still run whatever the file says.
GGUF and safetensors hold no Python, with one catch
A safetensors file is an 8-byte length, a JSON header of at most 100 MB, and raw numbers; a GGUF file is a typed list of metadata and tensor descriptions followed by aligned tensor data. Neither has a way to import or call anything. The catch in GGUF is the chat template: tokenizer.chat_template is a Jinja program that llama.cpp-based apps run on every prompt to lay out the conversation. In 2024 llama-cpp-python rendered it outside Jinja's sandbox (CVE-2024-34359), so a template that walked ''.__class__.__mro__ to Python's internals could run commands. This page shows the template in full and flags dunder attributes, OS and eval calls, and template imports in it.
A worked example: llama.cpp's own Llama vocabulary file
llama.cpp (MIT licence) keeps vocabulary-only GGUFs for its tokenizer tests. ggml-vocab-llama-spm.gguf is 723,869 bytes and opens as GGUF version 3 with 22 metadata keys and no tensors: architecture llama, context length 4,096, embedding length 4,096, 32 blocks with 32 attention heads, general.file_type 1 (F16), and a llama (SentencePiece) tokenizer whose tokenizer.ggml.tokens array holds 32,000 strings beginning <unk>, <s>, </s>, <0x00>. It has no chat template, and as a GGUF it imports nothing. The same reading on a 4 GB Q4_K_M model takes the first few megabytes of the file: the metadata and tensor list sit at the start, and the tensor data that follows is never read.
What this cannot do
- It does not run the model and cannot tell you how good it is: no inference, no accuracy check, no benchmark. It reads the file's structure, not what the numbers do.
- It cannot prove a file is safe. It shows what loading would import and call. A file with only safe-list imports cannot run a command through pickle, but its weights can still be backdoored or wrong, and a GGUF template check is a pattern match, not a Jinja sandbox.
- Not ONNX, Keras or TensorFlow. ONNX (.onnx), Keras .h5 and .keras files, and TensorFlow SavedModels are not read; the page says so when you drop one.
- Compressed joblib files are not read (zlib, gzip, xz or lz4, as
joblib.dump(..., compress=3)writes them). An uncompressed joblib file that stores NumPy arrays inline is walked up to the first array; the page says where it stopped, and an import after that point cannot be ruled out. - TorchScript code is listed, not judged. A TorchScript archive (from
torch.jit.save) holds Python-like source undercode/that runs as part of the model; the page names those files and their count. - The tensor values are not shown. Shapes, types and sizes are; the numbers themselves are not decoded.