Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
22 changes: 22 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
@@ -1,5 +1,27 @@
# Changelog

## 1.7.0 — unreleased

> **Unreadable model files now exit `1` instead of `0`.** A file that no parser could read — a git-LFS pointer stub, an empty file, a truncated stream, a file in no recognised format — was reported as `Risk: LOW` with a confident framework label and passed as a clean scan. It is now reported as an error and exits `1`. **If your tree contains such files, a pipeline that passes today will start failing.** That is the intended signal: those files were never examined, and `--fail-on-risk` was reading `LOW` on an artifact nothing had opened. The exit-code table in the README has always documented exit `1` for "a file failed to parse".

### Fixed

- **"Could not read it" is no longer reported as "read it, it's clean".** `LOW` asserts that AIsbom opened an artifact and found nothing dangerous; for a file no parser could read, the honest answer is different. Five inspectors already recorded their own parse failure internally and nothing ever consulted it, so a truncated zip, a two-byte pickle, 200 bytes of random data named `.safetensors`, and a plain-text `.pt` all graded `LOW` at exit `0`.

Such files now land in the scan's error list, print in their own **Could not read** section, and exit `1`. They still appear in the SBOM — an auditor needs to see the file was present — but carry no framework label and no risk verdict, marked with `aisbom:unreadable` and `aisbom:unreadable_type` properties. A format label is applied only when that format's parser actually succeeded, so random bytes are no longer described as SafeTensors. `aisbom score` consequently refuses to grade a scan containing one.

A **git-LFS pointer** gets its own message naming the remedy (`git lfs pull`), since that is a configuration problem rather than a corrupt artifact and is the most common unreadable `.safetensors` in practice.

- **`UNKNOWN` verdicts have a consequence.** An unparseable GGUF, Keras or ONNX file was already labelled honestly, but `UNKNOWN` scores below `LOW` in the risk ranking, so those scans exited `0` and the file rated *safer* than a clean model. They now exit `1` like every other unreadable file.

- **A text `.pt` or `.bin` is no longer classified as a Python path config.** `.pth` is legitimately also a Python path-configuration format, and any text file in those three extensions inherited that classification — so an HTML error page saved over a checkpoint scored `LOW`. The classification is now `.pth`-only and validates against that format's actual spec (`import` statements and bare paths), verified against every `.pth` file in a real virtualenv.

- **A SafeTensors header longer than the file is rejected before it is read.** Random bytes decode to an astronomical declared header length; that is now bounded by the bytes actually present, which replaces an internal `OverflowError` with a message saying what is wrong with the file.

### Changed

- **Telemetry reports unreadable files.** `cli_scan` gains `unreadable_count` and `unreadable_types` (a closed set: `LfsPointer`, `EmptyFile`, `TruncatedStream`, `UnrecognizedFormat`), so a wave of failed downloads is distinguishable from a wave of clones missing `git lfs pull`. Paths are never sent. `AISBOM_NO_TELEMETRY=1` still turns telemetry off.

## 1.6.0 — 2026-09-14

> **`--vex` now contacts a third-party service.** When a `--vex` scan finds exact `requirements.txt` pins, it sends each pinned package name and version to the public OSV API at `api.osv.dev`. Nothing else is sent. Pass `--no-osv` or set `AISBOM_NO_OSV=1` to turn this off. The GitHub Action runs `--vex` whenever its `token` input is set, so those runs make the lookup too.
Expand Down
16 changes: 16 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -134,6 +134,22 @@ xattr -d com.apple.quarantine aisbom-macos-*

`--no-fail-on-risk` governs risk findings only. An unusable target still exits `1`, so a typo'd path in CI fails loudly instead of passing as a clean scan.

### Files that could not be read

`LOW` is an assertion: AIsbom opened the artifact and found nothing dangerous. A file that no parser could read gets a different answer, reported in its own **Could not read** section and exiting `1`:

```
⚠️ Could not read:
- models/model.safetensors: this is a git-LFS pointer file, not the model
itself — run `git lfs pull` to fetch the real artifact, then re-scan
- models/head.pkl: declared a pickle by its extension, but nothing past the
protocol header could be disassembled — truncated or not a pickle at all
```

The cases are a git-LFS pointer stub (the usual one — a clone without `git lfs pull`), an empty file, a truncated stream, and a file in no recognised format. Such an artifact still appears in the SBOM, so an auditor can see it was present, but it carries **no framework label and no risk verdict** — a format label is only applied when that format's parser actually succeeded. In CycloneDX it is marked with `aisbom:unreadable` and `aisbom:unreadable_type` properties, and `aisbom score` refuses to grade a scan containing one.

> **Upgrading:** before v1.7.0 these files were reported as `Risk: LOW` with a confident framework label and exited `0`, so a truncated download could pass a CI gate as a clean model. If your tree contains unreadable files, those scans now exit `1`. That is the behaviour the exit-code table above always documented; the files were never examined.

### Scan a Hugging Face model

```bash
Expand Down
33 changes: 31 additions & 2 deletions aisbom/cli.py
Original file line number Diff line number Diff line change
Expand Up @@ -833,6 +833,10 @@ def _risk_score(label: str) -> int:
# (context) below. Intentional; see the slice notes.
fetch_failures = [e for e in results['errors'] if e.get('fetch_failure')]
target_errors = [e for e in results['errors'] if e.get('target_error')]
# Files a model extension claimed that no parser could read (#131). Split
# out here, ahead of telemetry, so the count and subtypes ride into
# `cli_scan` alongside the target-error split.
unreadable = [e for e in results['errors'] if e.get('unreadable')]
# Loop detection (#99) works at scan granularity ("N runs in a row"): the
# first failure's fingerprint represents the invocation — a fetch failure
# if there is one, else a target error (#126: a cron job on a typo'd path
Expand Down Expand Up @@ -896,6 +900,14 @@ def _risk_score(label: str) -> int:
# split that tells "scanned nothing" from "scanned a corrupt model".
"parse_error_count": str(len(results.get("errors", []))),
"target_error_count": str(len(target_errors)),
# #131, following the #126 split. The count answers "how much of this
# tree could we not read", and the closed-set subtypes say why — which
# is what tells a wave of failed downloads from a wave of clones
# missing `git lfs pull`. Both are bounded cardinality by construction.
"unreadable_count": str(len(unreadable)),
"unreadable_types": ",".join(sorted(
{e["unreadable_type"] for e in unreadable if e.get("unreadable_type")}
)),
"strict_mode": "true" if strict else "false",
}
telemetry_threads.append(
Expand Down Expand Up @@ -997,11 +1009,28 @@ def _risk_score(label: str) -> int:
if target_errors and not fetch_failures:
_maybe_print_loop_warning(loop_count, first_payload["http_status"])

# Parse errors only — fetch failures and target errors already printed
# their own message to stderr and don't fit the "Could not parse" framing.
# Files carrying a model extension that no parser could read (#131). Their
# own section, because "could not parse" describes a parser that ran and
# failed, while these mostly never got that far — an empty file, a git-LFS
# stub, a stream that stops mid-way. The distinction is the actionable
# part: the LFS case is a `git lfs pull` away from scanning fine.
if unreadable:
console.print("\n[bold red]⚠️ Could not read:[/bold red]")
for err in unreadable:
console.print(f" - [yellow]{err['file']}[/yellow]: {err['error']}")
console.print(
"\n[dim]These files were not examined, so they carry no risk "
"verdict. A scan that cannot read an artifact is not a scan that "
"found it clean.[/dim]"
)

# Parse errors only — fetch failures, target errors and unreadable files
# already printed their own message and don't fit the "Could not parse"
# framing.
parse_errors = [
e for e in results['errors']
if not e.get('fetch_failure') and not e.get('target_error')
and not e.get('unreadable')
]
if parse_errors:
console.print("\n[bold red]⚠️ Errors Encountered:[/bold red]")
Expand Down
12 changes: 12 additions & 0 deletions aisbom/properties.py
Original file line number Diff line number Diff line change
Expand Up @@ -66,6 +66,18 @@ def build_component_properties(art: Dict[str, Any]) -> List[Tuple[str, str]]:
if legal:
props.append(("aisbom:legal", str(legal)))

# A file no parser could read (#131). Emitted before the format branch
# because such an artifact deliberately carries no `framework` — the label
# is an assertion that that format's parser succeeded — so it returns below
# without any `aisbom:format`. The marker is what keeps the component
# honest rather than merely silent: a consumer sees the file was present
# and was not examined, instead of reading an absent risk as a clean one.
if art.get("unreadable"):
props.append(("aisbom:unreadable", "true"))
unreadable_type = art.get("unreadable_type")
if unreadable_type:
props.append(("aisbom:unreadable_type", str(unreadable_type)))

fmt = _format_for(art)
if fmt is None:
return props
Expand Down
30 changes: 30 additions & 0 deletions aisbom/safety.py
Original file line number Diff line number Diff line change
Expand Up @@ -1151,6 +1151,36 @@ def looks_like_pickle_stream(data: bytes) -> bool:
return False


def pickle_content_opcode_count(data: bytes) -> int:
"""How many opcodes past the protocol header ``data`` disassembles into.

Answers a narrower question than `looks_like_pickle_stream`: not "is this a
complete pickle" but "was there anything here to look at at all". Zero means
the disassembler got nothing beyond a bare `PROTO` byte — the file is
truncated or is not a pickle — which is what separates "could not read it"
from "read it, found nothing dangerous" (#131).

Reaching STOP is deliberately *not* required, because a real and complete
pickle need not be the whole file. joblib writes its arrays as raw bytes
directly after the pickle's STOP, so the opcode walk over a valid
`.pkl` from joblib's own test corpus dies on that trailing data after 54
content opcodes. Requiring STOP called 75 such files unreadable.

Disassembly only. The stream is never unpickled.
"""
count = 0
try:
for opcode, _arg, _pos in pickletools.genops(io.BytesIO(data)):
if opcode.name == "PROTO":
# The header says which protocol follows; it is not content.
continue
count += 1
except Exception:
# A walk that dies partway still examined everything it counted.
pass
return count


class _NullWriter:
"""Sink for `pickletools.dis`, which validates by writing a listing."""

Expand Down
Loading
Loading