From d297633124526a7bd35723ad3387b52d37308b53 Mon Sep 17 00:00:00 2001 From: BinFlip Date: Sat, 19 Sep 2026 06:23:47 -0700 Subject: [PATCH] chore: release 0.5.1 Maintenance release with no functional change. - Refresh transitive dependencies via cargo update. Direct dependencies (goblin 0.10.7, divan 0.1.21) are already current. - Replace em-dashes and en-dashes with plain hyphens across rustdoc, comments, the changelog and the sample-corpus docs. Three doc comments where a dash opened a line are rejoined to the preceding line, since a leading "- " renders as a Markdown list item and trips clippy::doc_lazy_continuation. - Add the missing 0.5.0 compare link to the changelog. --- CHANGELOG.md | 193 +++++++++++++++++++---------------- Cargo.lock | 53 ++++++---- Cargo.toml | 4 +- benches/extract.rs | 12 +-- src/detection.rs | 8 +- src/formats.rs | 50 ++++----- src/lib.rs | 64 ++++++------ src/metadata.rs | 46 ++++----- src/structures/abitype.rs | 6 +- src/structures/buildinfo.rs | 10 +- src/structures/embed.rs | 6 +- src/structures/fixups.rs | 8 +- src/structures/gcprog.rs | 8 +- src/structures/gostring.rs | 2 +- src/structures/inittask.rs | 4 +- src/structures/inline.rs | 12 +-- src/structures/itab.rs | 10 +- src/structures/locate.rs | 10 +- src/structures/maptype.rs | 34 +++--- src/structures/mod.rs | 14 +-- src/structures/moduledata.rs | 70 ++++++------- src/structures/pclntab.rs | 56 +++++----- src/structures/strings.rs | 14 +-- src/structures/types.rs | 46 ++++----- src/structures/util.rs | 2 +- src/structures/wasm.rs | 14 +-- tests/integration.rs | 58 +++++------ tests/samples/README.md | 30 +++--- tests/samples/build.sh | 10 +- 29 files changed, 439 insertions(+), 415 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 13093e8..d73d81b 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -5,6 +5,17 @@ All notable changes to `gobin` are documented in this file. The format is based on [Keep a Changelog](https://keepachangelog.com/en/1.1.0/), and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0.html). +## [0.5.1] + +Maintenance release. No functional change. + +### Changed + +- Refreshed transitive dependencies (`cargo update`); no direct dependency + changed version (`goblin` 0.10.7 and `divan` 0.1.21 are current). +- Replaced em-dashes and en-dashes with plain hyphens throughout rustdoc, + comments, the changelog and the sample-corpus docs. + ## [0.5.0] Go 1.27 verified against the released toolchain, plus a performance pass. @@ -15,7 +26,7 @@ so what needed fixing was the code paths it shares with older versions. ### Fixed - **The V5 layout could be chosen for a pre-1.27 binary, silently emptying it.** - It was selected from a *negative* signal — no `.typelink` section — which PE + It was selected from a *negative* signal - no `.typelink` section - which PE never has at any version, nor wasm, nor RELRO ELF. With no version string to disagree, a Go 1.20-1.26 binary decoded as V5 and every field past `types` shifted: on a version-scrubbed `basic_go124_windows_amd64.exe`, types went @@ -28,7 +39,7 @@ so what needed fixing was the code paths it shares with older versions. bytes on 1.24-1.26, 48 on 1.14-1.23), so every `TypeDetail::Map` field past `Hasher` was garbage, and `descriptor_size` mislocated the trailing `UncommonType`. `MapLayout` now models all six eras, selected from the Go - version or — when scrubbed — from the moduledata version and pclntab magic, + version or - when scrubbed - from the moduledata version and pclntab magic, with a structural probe for the one window those cannot separate. `MapFlags` normalizes the three flag encodings. - **`text_va()` returned `Some(0)` on Go 1.26+ without a moduledata.** Go 1.26 @@ -36,7 +47,7 @@ so what needed fixing was the code paths it shares with older versions. `entry_va` collapsed to a raw `entry_off`. Zero now reads as absent, and the executable section base is used as a real fallback. - **The Go 1.27 type walk ran past its bound**, using `etypes` where - `runtime.moduleTypelinks` uses `types + typedesclen` — a field that was + `runtime.moduleTypelinks` uses `types + typedesclen` - a field that was parsed but never read. Past the bound it surfaced unnamed junk types. - **One unparseable record ended the whole descriptor walk**, which on Go 1.27 is the only type-enumeration strategy there is. It now skips and continues @@ -63,7 +74,7 @@ so what needed fixing was the code paths it shares with older versions. profiling and stage-level benches (`stage_context`, `stage_buildinfo`, `stage_pclntab`). -### Changed — breaking +### Changed - breaking - `Moduledata::parse` takes a `LayoutHints` struct instead of a `(PclntabVersion, bool, Option)` tail. @@ -93,8 +104,8 @@ Extraction is 47-74% faster and allocates about a fifth as much. `full_sweep` - The build-info magic scan walked the image twice, the "aligned" pass one byte at a time. Now one sweep, narrowed first to the regions that can hold a `SBUILDINFO` symbol: PE `parse` 4.10 ms → **160 µs**, wasm → 1.52 ms. -- `types()` / `all_types()` / `type_at()` re-ran moduledata discovery — scan - included — on every call; they now reuse the parsed one: **−88% PE, −90% wasm**. +- `types()` / `all_types()` / `type_at()` re-ran moduledata discovery - scan + included - on every call; they now reuse the parsed one: **−88% PE, −90% wasm**. - `init_order()` indexed all ~1800 functions to resolve a dozen init PCs: **1.12 ms → 22 µs, 206 KB → 2.1 KB**. - `embedded_assets()` allocated two `String`s per candidate entry for a @@ -137,21 +148,21 @@ Extraction is 47-74% faster and allocates about a fifth as much. `full_sweep` ### Added -Support for the legacy Go 1.2–1.15 binary layout, which previously parsed as +Support for the legacy Go 1.2-1.15 binary layout, which previously parsed as garbage (a Go 1.9.x binary produced millions of bogus source-file names and unprintable "package" strings). The pre-1.16 pclntab and moduledata are genuinely different structures, now parsed natively: - `pclntab::parse_header_go12` handles the legacy "Go 1.2" pclntab (magic `0xfffffffb`): an 8-byte header (no structured `pcHeader`), a pointer-sized - functab of absolute PCs, and the `[]uint32` `filetab` — the table that + functab of absolute PCs, and the `[]uint32` `filetab` - the table that previously misparsed into millions of entries. Function names and the `pcsp`/`pcfile`/`pcln` tables are pcHeader-relative; `text_va` is recovered from the lowest function PC when no moduledata is present. -- `Moduledata::parse_go12_legacy` parses the V1 moduledata (Go 1.5–1.15) with +- `Moduledata::parse_go12_legacy` parses the V1 moduledata (Go 1.5-1.15) with per-minor field gating verified against `runtime/symtab.go`: `itablinks` (1.6), `types`/`typelinks []int32`/`typemap` (1.7), plugin fields (1.8), - `hasmain`/`bad` (1.10). Go 1.2–1.4 have no moduledata and are left as such. + `hasmain`/`bad` (1.10). Go 1.2-1.4 have no moduledata and are left as such. The moduledata locators (in `lib` and `types`) validate the legacy layout through its `text` boundary, since it has no `funcnametab`. @@ -172,7 +183,7 @@ parameter array, but the parser skipped only the bare `FuncTypeExtra` (4 bytes), landing the `UncommonType` short by the padding. That read a garbage `mcount` (up to `0xFFFF` from padding bytes), and resolving those phantom methods followed garbage `mtyp` offsets into unrelated -descriptors — inflating `all_types()` from the ~850 genuinely-reachable +descriptors - inflating `all_types()` from the ~850 genuinely-reachable types to ~12.8k and surfacing stdlib structs not actually reachable from typelinks. @@ -180,7 +191,7 @@ typelinks. `Func` extra (mirroring `descriptor::descriptor_size`), so the `UncommonType` is located correctly. - `read_func_params` skips the `UncommonType` block when present, since Go - places the inline parameter array *after* it (`abi.FuncType.InSlice`) — + places the inline parameter array *after* it (`abi.FuncType.InSlice`) - keeping the parameter VAs (and the types they reach) correct. - `resolve_concrete_methods` / `resolve_interface_methods` now stop on the first empty/unresolved method name rather than fabricating thousands of @@ -195,14 +206,14 @@ and committed sample fixtures (see `tests/samples/README.md`). ### Added -Go 1.27 ("V5") moduledata support — the upstream runtime removed the +Go 1.27 ("V5") moduledata support - the upstream runtime removed the `typelinks`/`itablinks` slices and now stores interface tables inline: - `ModuledataVersion::V5` is fully parsed: `itaboffset`/`itabsize`, `typedesclen`, and the still-present `rodata`/`gofunc`/`inittasks` fields. The previous speculative V5 branch stopped at `etypes` and hardcoded `rodata`/`gofunc` to `None`, which **silently disabled** `inline_tree()` - and `itab_pairs()` on Go 1.27 binaries — now fixed. + and `itab_pairs()` on Go 1.27 binaries - now fixed. - `itab_pairs()` gained a third strategy: it walks the inline, variable-size itab records at `types + itaboffset` (Go 1.27+), in addition to the `.itablink` section and `moduledata.itablinks` slice used by Go ≤1.26. @@ -211,20 +222,20 @@ Go 1.27 ("V5") moduledata support — the upstream runtime removed the New extraction surfaces: -- `bin.fips_info() -> Option` — FIPS-140 mode (the `GOFIPS140` +- `bin.fips_info() -> Option` - FIPS-140 mode (the `GOFIPS140` build setting) plus the `__go_fipsinfo` integrity sum. Returns `None` for non-FIPS builds (the integrity section is linked into every Go 1.24+ binary, so it alone is not a FIPS signal). -- `bin.init_order() -> Vec` — package initialization order decoded +- `bin.init_order() -> Vec` - package initialization order decoded from `moduledata.inittasks` (Go 1.24+), with init-function entry VAs resolved to names. Captured for both V4 and V5 layouts. -- `bin.embedded_assets() -> Vec` — `//go:embed` payloads +- `bin.embedded_assets() -> Vec` - `//go:embed` payloads (path, dir flag, backing bytes) decoded from `embed.FS` `[]file` arrays. Symbol-independent (works on stripped binaries) and cross-format (ELF/Mach-O/PE/wasm); validated by enforcing `embed`'s canonical sort order and dir/file hash consistency to avoid false positives. The single-file `embed.String`/`embed.Bytes` form is not recovered (no anchor). -- `BuildInfo::deps_full()` and `BuildInfo::module_sum()` — surface the +- `BuildInfo::deps_full()` and `BuildInfo::module_sum()` - surface the already-parsed `go.sum` hashes and `replace` directives that the `dependencies()` convenience iterator collapses (supply-chain analysis). @@ -252,34 +263,34 @@ verified against go1.16-1.27 `runtime/symtab.go`: (equality-function) and `GCData` (GC-bitmap) pointers it previously skipped, and `GoType` exposes the full descriptor: `descriptor_va` (its own VA), `tflag` (raw flag byte), `ptr_to_this` (the `*T` `TypeOff`), `equal_va`, `gcdata_va`, -and `pkg_path` — the import path resolved from the `UncommonType` (distinct +and `pkg_path` - the import path resolved from the `UncommonType` (distinct from the name-derived `package()`). The kind-specific `TypeDetail` variants and `MethodEntry` now carry every parsed field: `Array.slice_va`; `Map.group_va`/`hasher_va`/`key_stride`/ `elem_stride`/`flags`; `Interface.pkg_path`; struct `StructField.tag` (the decoded struct tag, e.g. `json:"id"`); and `MethodEntry.interface_text_offset` -(the `ifn` wrapper entry, alongside the `tfn` direct entry — both now correctly +(the `ifn` wrapper entry, alongside the `tfn` direct entry - both now correctly treat the `-1` "no body" sentinel as `None`). **Deep type / runtime enumeration.** New accessors: -- `bin.all_types()` — transitively enumerates **every** reachable type +- `bin.all_types()` - transitively enumerates **every** reachable type descriptor (BFS from the `typelink` seed set, following element / key / field / parameter / method / pointer-to-this references, each parsed independently by VA and bounded to `[types, etypes)`). Reaches types absent from `typelink` (e.g. a struct used only as a pointer's element, with its field tags). `bin.type_at(va)` parses a single descriptor by address. -- `bin.data_pointer_map()` / `bss_pointer_map()` — decode the `gcdata` / +- `bin.data_pointer_map()` / `bss_pointer_map()` - decode the `gcdata` / `gcbss` **GC programs** (`runGCProg` bytecode) into a per-word pointer bitmap: the precise location of every pointer in global memory (function pointers, `itab`/interface pointers, string/slice headers, global `*T`), recoverable without disassembly. New `gcprog::{run_gc_prog, PointerMap}`. -- `bin.modules()` — walk the `moduledata.next` linked list (one module for a +- `bin.modules()` - walk the `moduledata.next` linked list (one module for a normal static binary; defensive for multi-module images). -- `bin.itab_methods(pair)` — the `Fun[]` method-pointer array of an itab. +- `bin.itab_methods(pair)` - the `Fun[]` method-pointer array of an itab. - `bin.text_sections()`, `bin.plugin_exports()` (ptab), `bin.package_hashes()` - / `bin.module_hashes()` (per-package link-time ABI hashes) — decode the + / `bin.module_hashes()` (per-package link-time ABI hashes) - decode the remaining moduledata sub-slices (`textsectmap`, `ptab`, `pkghashes`, `modulehashes`). - `ParsedPclntab` gained `header_text_start` (the pcHeader `textStart` field, @@ -291,8 +302,8 @@ treat the `-1` "no body" sentinel as `None`). `-buildmode=plugin` / c-shared) store data pointers as `LC_DYLD_CHAINED_FIXUPS` chains rather than absolute values. The crate now walks those chains and rebases every pointer into an owned shadow image - (`structures::macho_fixups`), so pointer-dependent extraction — types, - itabs, `ptab` plugin exports, per-package hashes, plugin path — works on + (`structures::macho_fixups`), so pointer-dependent extraction - types, + itabs, `ptab` plugin exports, per-package hashes, plugin path - works on these objects exactly as on a pure-Go executable. Supports the `DYLD_CHAINED_PTR_64` and `DYLD_CHAINED_PTR_64_OFFSET` pointer formats. Pure-Go executables are unaffected (no fixups load command → no rebasing). @@ -309,7 +320,7 @@ The `dump` example prints a new "Runtime & Linker Surfaces" section. display name). `StructField`, `MethodEntry`, and `InterfaceMethod` each gained a `type_name: Option<&str>` field. Method/interface-method signature types are unnamed Go func types, so their `type_name` is usually - `None` — resolve the offset to a `TypeDetail::Func` for the full, + `None` - resolve the offset to a `TypeDetail::Func` for the full, name-resolved signature. Leaf (param/field) types resolve to names. - `itab::extract_iter` now takes `Option<&Moduledata>` instead of `Option<&GoSlice>` (it needs `itaboffset`/`itabsize` for V5). @@ -317,13 +328,13 @@ The `dump` example prints a new "Runtime & Linker Surfaces" section. parser (`elf::Elf`, `mach::MachO`, `pe::PE`) instead of the unified `goblin::Object::parse`, dropping the `te` (UEFI Terse Executable) default feature. PE parsing runs with `ParseOptions` that disable resource, import, - certificate, and TLS parsing — none of which the crate reads — avoiding + certificate, and TLS parsing - none of which the crate reads - avoiding wasted work on Go binaries with large import/resource tables. No change to extracted data. ### Fixed -**Go 1.16-1.17 support** — the parser previously handled only Go 1.18+. The +**Go 1.16-1.17 support** - the parser previously handled only Go 1.18+. The pre-1.18 binary format is now decoded: the pcHeader without `textStart`, the `2×ptrSize` functab with absolute PCs, the `_func` struct led by an absolute `entry uintptr`, and the pre-1.17 type-name encoding (2-byte big-endian length @@ -341,7 +352,7 @@ against real go1.16 / 1.19 / 1.21 / 1.24 / 1.26 / 1.27 toolchains: - **`rodata`/`gofunc` are gated to Go 1.18+**, separately from `covctrs` (Go 1.20+). They were previously gated together on the Go120 magic, so - 1.18-1.19 binaries read `typelinks`/`itablinks` two pointers early — breaking + 1.18-1.19 binaries read `typelinks`/`itablinks` two pointers early - breaking type/itab extraction on PE 1.18-1.19 (ELF/Mach-O were masked by the section-based path). Moduledata `V2` now means 1.16-1.17; `V3` is 1.18-1.25. @@ -371,7 +382,7 @@ upgrades, and a redesigned pclntab caching model that does not leak memory. ### Added -WebAssembly support — `GOOS=js GOARCH=wasm` (and the newer `wasip1` Go 1.21+ +WebAssembly support - `GOOS=js GOARCH=wasm` (and the newer `wasip1` Go 1.21+ target) now parses end-to-end with the same surface as ELF / Mach-O / PE: - New `BinaryFormat::Wasm` variant; `\0asm` magic detection in @@ -392,11 +403,11 @@ target) now parses end-to-end with the same surface as ELF / Mach-O / PE: New accessors and ergonomics helpers: -- `bin.entry_va(func)` / `bin.entry_rva(func)` — fold +- `bin.entry_va(func)` / `bin.entry_rva(func)` - fold `text_va + entry_off` (and the PE-only `image_base` subtraction) into one accessor so downstream disassemblers do not have to reimplement it. `BinaryContext::image_base()` exposed in support. -- `bin.arch()` — format-disambiguated architecture accessor that resolves +- `bin.arch()` - format-disambiguated architecture accessor that resolves the `(minLC=1, ptrSize=8)` ambiguity between `Arch::X86_64` and `Arch::Wasm` using the container format, and falls back to build-info `GOARCH` when no pclntab is present. @@ -406,7 +417,7 @@ New accessors and ergonomics helpers: on each type. - `BuildMode::as_str() -> Cow<'static, str>` + `Display`, round-tripping through `BuildMode::parse`. -- `ParsedPclntab::decode_pcln_with_files(func)` — joined iterator yielding +- `ParsedPclntab::decode_pcln_with_files(func)` - joined iterator yielding `(pc, line, file_path)` per source-line transition with the active file pre-attached, replacing the "walk pcln and pcfile in lockstep" loop callers otherwise reinvent. @@ -416,7 +427,7 @@ New accessors and ergonomics helpers: `is_stdlib`) plus shared free helpers `is_runtime_path` / `is_internal_path` / `is_stdlib_path` so type-side and function-side classifications agree on one canonical rule. -- `FuncData::args_size() -> u32` — typed view of the raw signed `args` +- `FuncData::args_size() -> u32` - typed view of the raw signed `args` field (negative readings collapse to `0`). - `GoString::as_bytes()` and `try_as_str() -> Result<&str, Utf8Error>`. The string scanner no longer silently drops non-UTF-8 length-prefixed @@ -427,9 +438,9 @@ Address-space + caching internals: - `BinaryContext::structure_search_data()` returns the wasm linear-memory image for wasm and the file bytes otherwise; `BinaryContext::va_to_file` produces offsets into this view for every format. -- `BinaryContext::slice_at_va(va)` — borrow a slice starting at the byte +- `BinaryContext::slice_at_va(va)` - borrow a slice starting at the byte at virtual address `va`; spans wasm data-segment boundaries seamlessly. -- `PclntabMeta` — scalar metadata of a `ParsedPclntab` without the +- `PclntabMeta` - scalar metadata of a `ParsedPclntab` without the borrowed `data`. Cache it; rehydrate via `PclntabMeta::attach(data)`. `ParsedPclntab::meta()` extracts it. - `ParsedPclntab` is now `Copy` (every field already was). @@ -439,23 +450,23 @@ Address-space + caching internals: `GoBinary` no longer caches a `ParsedPclntab` borrowing from the input. Instead it caches the scalar `PclntabMeta` and rebuilds the borrowing struct against `&self.ctx.structure_search_data()` per call. This is what -allows wasm support without a `Box::leak` — the wasm linear-memory image +allows wasm support without a `Box::leak` - the wasm linear-memory image is owned by `BinaryContext` and borrows handed out to callers tie to `&self`. -- `GoBinary::pclntab(&self) -> Option>` — was +- `GoBinary::pclntab(&self) -> Option>` - was `Option<&ParsedPclntab<'a>>` (breaking). Method calls and `let Some(pcl) = ...` patterns are unchanged; sites that explicitly named the reference type must drop the `&`. -- `FunctionIter::new(pcl: Option>)` — takes the struct +- `FunctionIter::new(pcl: Option>)` - takes the struct by value (it's `Copy`) instead of by reference. The iterator is no longer parameterized on the pcl-borrow lifetime: `FunctionIter<'a>` (was `FunctionIter<'p, 'a>`). -- `TypeIter`, `ItabIter`, `GoStringIter`, `InlineTreeIter` — lost their +- `TypeIter`, `ItabIter`, `GoStringIter`, `InlineTreeIter` - lost their `'ctx` / `'pcl` lifetime parameter; yielded items borrow from `&BinaryContext` (was the input lifetime). Callers that collect into a `Vec` keep working as long as they hold the `GoBinary`. -- `pclntab::parse(ctx)` now takes `&'a BinaryContext<'a>` — required so +- `pclntab::parse(ctx)` now takes `&'a BinaryContext<'a>` - required so the returned struct can borrow from either the input bytes or the wasm linear-memory image with one lifetime. - Stability-policy sections lifted to type-level rustdoc on `Confidence`, @@ -466,7 +477,7 @@ is owned by `BinaryContext` and borrows handed out to callers tie to `as_str` for text or `as_bytes` for raw rodata. - `ParsedPclntab::arch()` rustdoc clarified: `(1, 8)` always returns `Arch::X86_64`; use `GoBinary::arch()` for format-aware disambiguation. -- `text_va()` rustdoc no longer publishes a manual translation recipe — +- `text_va()` rustdoc no longer publishes a manual translation recipe - it points to the new `entry_va` / `entry_rva` instead. - Lints moved from a `#![deny(...)]` attribute on `lib.rs` to a `[lints]` table in `Cargo.toml`, so they enforce on every consuming workspace @@ -502,117 +513,117 @@ is owned by `BinaryContext` and borrows handed out to callers tie to ## [0.2.0] A large feature pass plus internal hardening for use in malware analysis -pipelines. **Many breaking API changes** — see *Removed* and *Changed*. +pipelines. **Many breaking API changes** - see *Removed* and *Changed*. -### Added — new extraction surfaces +### Added - new extraction surfaces -- **Per-PC inlining tree** — `bin.inline_tree(func)` yields +- **Per-PC inlining tree** - `bin.inline_tree(func)` yields `inline::InlineEntry { pc_range, function_name, parent_pc, start_line, func_id, depth }` per PC range with cycle-safe parent-chain walk for depth computation. -- **Go string literal scanner** — `bin.strings()` yields `GoString<'a> { +- **Go string literal scanner** - `bin.strings()` yields `GoString<'a> { va, len, bytes }` for every `(ptr, len)` header that resolves to in-binary UTF-8. Recovers strings a generic byte-string extractor would miss or split on internal NULs. -- **Itab pairs** — `bin.itab_pairs()` yields `ItabPair { iface_type_va, +- **Itab pairs** - `bin.itab_pairs()` yields `ItabPair { iface_type_va, concrete_type_va, hash, itab_va }` for every `(interface, concrete type)` pair the linker proved at build time. -- **Per-function inlining accessors** — `FuncData::func_off`, +- **Per-function inlining accessors** - `FuncData::func_off`, `ParsedPclntab::pcdata_at`, `ParsedPclntab::funcdata_at` for the variable-length tables after the `_func` 44-byte prefix. -- **Garble obfuscation detection** — `bin.obfuscation()` returns +- **Garble obfuscation detection** - `bin.obfuscation()` returns `ObfuscationKind { None, Garble { confidence }, Other { reason } }`; `bin.is_likely_garbled()` convenience. -- **Compiler identification** — `bin.compiler()` returns `Compiler { Gc, +- **Compiler identification** - `bin.compiler()` returns `Compiler { Gc, TinyGo, Gccgo, Unknown }`. -- **Cgo / concurrency presence** — `bin.has_cgo()` and +- **Cgo / concurrency presence** - `bin.has_cgo()` and `bin.uses_concurrency()` short-circuit on the first matching function. Per-call-site enumeration deferred (needs disassembler). -- **Runtime address accessors** — `bin.text_va()`, `bin.etext_va()` expose +- **Runtime address accessors** - `bin.text_va()`, `bin.etext_va()` expose `runtime.text` / `runtime.etext` with documented translation recipe for `entry_off → VA / RVA`. -- **Runtime commit hash** — `bin.runtime_commit()` extracts the dev commit +- **Runtime commit hash** - `bin.runtime_commit()` extracts the dev commit from `devel go1.X-` version strings. -- **Build mode / tags / dependencies** — `BuildInfo::build_mode()` returns +- **Build mode / tags / dependencies** - `BuildInfo::build_mode()` returns `BuildMode` enum; `build_tags()` iterates `-tags`; `dependencies()` and `build_settings_iter()` provide iterator accessors. -- **Module replacements + sums** — `DepEntry { path, version, sum, +- **Module replacements + sums** - `DepEntry { path, version, sum, replacement }` + `DepReplacement` parsed from modinfo `dep` / `=>` / sum lines. -- **Method extraction on types** — `GoType.methods: Vec` for +- **Method extraction on types** - `GoType.methods: Vec` for every type with an `UncommonType`. Resolves names, type-descriptor offsets, text offsets, and exported flag. -- **Deep type structure** — `TypeDetail` extended with: +- **Deep type structure** - `TypeDetail` extended with: - `Struct.fields: Vec` with name / type VA / offset / embedded - `Interface.methods: Vec` - `Map { key_va, elem_va }`, `Pointer { elem_va }`, `Slice { elem_va }`, `Chan { dir, elem_va }`, `Array { len, elem_va }` - `Func { in_count, out_count, is_variadic, inputs: Vec, outputs: Vec }` -- **`FuncFlags` newtype** — `FunctionInfo::func_flags()` returns typed view +- **`FuncFlags` newtype** - `FunctionInfo::func_flags()` returns typed view of the `_func.flag` byte; `is_top_frame()`, `is_sp_write()`, `is_asm()`, `is_systemstack()` accessors. -- **Receiver parsing** — `FunctionInfo::receiver_type() -> +- **Receiver parsing** - `FunctionInfo::receiver_type() -> Option` plus `method_name()` and `generic_args()` accessors. -- **Per-PC file resolution** — `ParsedPclntab::resolve_file_via_cu` is now +- **Per-PC file resolution** - `ParsedPclntab::resolve_file_via_cu` is now `pub`; `decode_pcfile_paths(func)` streams `(pc, &str)` per inlined region. -- **Structured detection report** — `GoBinary::try_parse() -> +- **Structured detection report** - `GoBinary::try_parse() -> Result` returning `ConfidenceReport` of typed `ConfidenceSignal` variants on success or failure. `bin.report()` accessor exposes the same on success. -- **Bulk function decoder** — `metadata::for_each_function(pcl, |info, +- **Bulk function decoder** - `metadata::for_each_function(pcl, |info, tables|)` walks every function with reusable per-PC table buffers, amortizing allocation across the whole binary. -- **Fast detection** — `gobin::detect(&[u8]) -> bool` does magic-byte + +- **Fast detection** - `gobin::detect(&[u8]) -> bool` does magic-byte + buildinfo header check without invoking `goblin` parse. -- **`examples/dump --explain`** — prints structured detection report +- **`examples/dump --explain`** - prints structured detection report (Confidence tier + per-signal breakdown). -- **moduledata accessors** — `bin.moduledata()`, `Moduledata::rodata`, +- **moduledata accessors** - `bin.moduledata()`, `Moduledata::rodata`, `Moduledata::gofunc` exposed. -- **Iterator-style API throughout** — `bin.functions()`, `bin.types()`, +- **Iterator-style API throughout** - `bin.functions()`, `bin.types()`, `bin.itab_pairs()`, `bin.strings()`, `bin.inline_tree()` are all true streaming iterators (`FunctionIter`, `TypeIter`, `ItabIter`, `GoStringIter`, `InlineTreeIter`). Per-PC table decoders likewise (`PcValueIter`, `PcLineIter`, `PcFileIter`, `PcFilePathIter`). -- **Property tests** — `package() + "." + short_name() == name` round-trip +- **Property tests** - `package() + "." + short_name() == name` round-trip property test plus a corpus of well-known Go function-name shapes. - **Centralized helpers** in `structures::util`: `slice_at::`, `advance`, `advance_n`, `align_up`, `align_up_u64`, `read_uvarint`, - `read_uintptr`, `read_u32`, `read_i32`, `read_u16` — single source of + `read_uintptr`, `read_u32`, `read_i32`, `read_u16` - single source of truth for offset arithmetic and primitive reads. -### Changed — borrowed metadata types (breaking) +### Changed - borrowed metadata types (breaking) All metadata types now borrow from the input binary via lifetime `'a`, matching `FunctionInfo<'a>`. Callers that need to outlive the binary's lifetime must `.to_owned()` at the boundary. -- `GoType` → `GoType<'a>` — `name: String` becomes `&'a str`; `methods: +- `GoType` → `GoType<'a>` - `name: String` becomes `&'a str`; `methods: Vec>`; `detail: TypeDetail<'a>`. -- `MethodEntry`, `StructField`, `InterfaceMethod` — same treatment. -- `BuildInfo` → `BuildInfo<'a>` — all string fields borrow from the modinfo +- `MethodEntry`, `StructField`, `InterfaceMethod` - same treatment. +- `BuildInfo` → `BuildInfo<'a>` - all string fields borrow from the modinfo blob. - `DepEntry` → `DepEntry<'a>`, `DepReplacement` → `DepReplacement<'a>`. -- `GoBinary.go_version` and `GoBinary.build_id` — now `Option<&'a str>` +- `GoBinary.go_version` and `GoBinary.build_id` - now `Option<&'a str>` storage; accessors return `Option<&'a str>`. -### Changed — name parsing rewrites (breaking semantics) +### Changed - name parsing rewrites (breaking semantics) -- `FunctionInfo::package()` and `short_name()` rewritten — boundary is now +- `FunctionInfo::package()` and `short_name()` rewritten - boundary is now "first `.` after the last `/`" plus a `gopkg.in`-style `.vN` extension. Third-party functions like `github.com/spf13/cobra.(*Command).Run` now return `package = "github.com/spf13/cobra"` instead of `"github"`. -- `FunctionInfo::is_method()` — structural parser. Catches value-receiver +- `FunctionInfo::is_method()` - structural parser. Catches value-receiver methods (`time.Time.String`) the old `".("` substring heuristic missed, and excludes closures. -- `FunctionInfo::is_closure()` — strict: requires `.funcN` / `.gowrapN` +- `FunctionInfo::is_closure()` - strict: requires `.funcN` / `.gowrapN` numeric suffix and excludes asm-flagged functions. -- `decode_pcvalue` for pcfile — `decode_pcfile(func)` yields `(u32, u32)` +- `decode_pcvalue` for pcfile - `decode_pcfile(func)` yields `(u32, u32)` (was `(u32, i32)`); file indices are unsigned. -### Changed — streaming-only iterator API (breaking) +### Changed - streaming-only iterator API (breaking) Every `Vec`-returning method that had a streaming counterpart was dropped. Iterators replace them under the same names: @@ -630,22 +641,22 @@ Iterators replace them under the same names: To get an owned `Vec`, call `.collect()`. -### Changed — module renames +### Changed - module renames - `structures::gostring::GoString` → `GoStringHeader` (it was always just the `(ptr, len)` header pair). The `GoString<'a>` name now belongs to the public scanned-string type in `structures::strings`. -### Removed — legacy convenience APIs (breaking) +### Removed - legacy convenience APIs (breaking) -- `metadata::extract_functions(pcl) -> Vec` — use +- `metadata::extract_functions(pcl) -> Vec` - use `bin.functions()` or `FunctionIter::new(Some(pcl))`. -- `bin.types_iter()` / `bin.itab_pairs_iter()` aliases — the canonical +- `bin.types_iter()` / `bin.itab_pairs_iter()` aliases - the canonical names (`bin.types()` / `bin.itab_pairs()`) now return the iterators. - `pclntab.decode_pcvalue_into / decode_pcln_into / decode_pcfile_into` - buffer-reuse variants — use `buf.clear(); buf.extend(decode_*(...))` + buffer-reuse variants - use `buf.clear(); buf.extend(decode_*(...))` with the streaming iterators (same allocation behavior). -- `ParsedPclntab::read_ptr` — was dead code internally. +- `ParsedPclntab::read_ptr` - was dead code internally. ### Fixed @@ -658,7 +669,7 @@ To get an owned `Vec`, call `.collect()`. - `gopkg.in/yaml.v3.Marshal`-style names no longer split as `package = "gopkg.in/yaml"` (now `"gopkg.in/yaml.v3"`). -### Security — panic-free lint sweep +### Security - panic-free lint sweep This crate is used for malware analysis: every input byte is adversarial and must not be allowed to panic the parser. @@ -694,9 +705,9 @@ any `parse` / `extract` / `decode` path. Failures degrade to `None` / functions with inlining, 24,438 entries, depth distribution `{0: 17686, 1: 5604, 2: 1022, 3: 124, 4: 2}`. - String scanner verified across Mach-O / ELF / PE plus stripped - binaries (1k–1.5k unique strings recovered per binary). + binaries (1k-1.5k unique strings recovered per binary). -## [0.1.0] — initial release +## [0.1.0] - initial release Initial public release. @@ -707,6 +718,8 @@ Initial public release. - Type descriptor extraction via `.typelink` and descriptor walking. - Heuristic confidence scoring (`Confidence` enum). +[0.5.1]: https://github.com/ATRAPSLLC/gobin/compare/v0.5.0...v0.5.1 +[0.5.0]: https://github.com/ATRAPSLLC/gobin/compare/v0.4.1...v0.5.0 [0.4.1]: https://github.com/ATRAPSLLC/gobin/compare/v0.4.0...v0.4.1 [0.4.0]: https://github.com/ATRAPSLLC/gobin/compare/v0.3.1...v0.4.0 [0.3.1]: https://github.com/ATRAPSLLC/gobin/compare/v0.3.0...v0.3.1 diff --git a/Cargo.lock b/Cargo.lock index d11ab85..17ac156 100644 --- a/Cargo.lock +++ b/Cargo.lock @@ -10,30 +10,30 @@ checksum = "940b3a0ca603d1eade50a4846a2afffd5ef57a9feac2c0e2ec2e14f9ead76000" [[package]] name = "bitflags" -version = "2.13.1" +version = "2.13.2" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "b588b76d00fde79687d7646a9b5bdf3cc0f655e0bbd080335a95d7e96f3587da" +checksum = "3ded4057c258ba199e2d26386d3af3780957ecaee6c4ef4041c6b4b8b97c0b06" [[package]] name = "cfg-if" -version = "1.0.4" +version = "1.0.5" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "9330f8b2ff13f34540b44e946ef35111825727b38d33286ef986142615121801" +checksum = "4e7648175b45a9a48536d676f68d918270699102aa8dab5496df06904c914600" [[package]] name = "clap" -version = "4.6.6" +version = "4.6.7" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "473c7e07f409a8d772161724aa8db6a765a2532a70f9667eeb7b49d3d02fbdca" +checksum = "aa8876b300ab35ba921adea3dfd70157a46249b33f95c9084ae5709785478946" dependencies = [ "clap_builder", ] [[package]] name = "clap_builder" -version = "4.6.6" +version = "4.6.7" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "7b48fea5a88e9ae728a2dcbedbfc0e730f7d60da42e1cb049a83c9fb8b789889" +checksum = "ec0797fb7aeb1406c84efac526901f7ec3ead2124f946b494e72879d4b54704d" dependencies = [ "anstyle", "clap_lex", @@ -42,9 +42,9 @@ dependencies = [ [[package]] name = "clap_lex" -version = "1.1.0" +version = "1.1.1" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "c8d4a3bb8b1e0c1050499d1815f5ab16d04f0959b233085fb31653fbfc9d98f9" +checksum = "1c133bc6a41be0d194c306b5506d15e6feeea7b1d6604bd3f8310dfb2ca96486" [[package]] name = "condtype" @@ -74,7 +74,7 @@ checksum = "9556bc800956545d6420a640173e5ba7dfa82f38d3ea5a167eb555bc69ac3323" dependencies = [ "proc-macro2", "quote", - "syn", + "syn 2.0.119", ] [[package]] @@ -89,7 +89,7 @@ dependencies = [ [[package]] name = "gobin" -version = "0.5.0" +version = "0.5.1" dependencies = [ "divan", "goblin", @@ -120,9 +120,9 @@ checksum = "32a66949e030da00e8c7d4434b251670a91556f4144941d37452769c25d58a53" [[package]] name = "log" -version = "0.4.33" +version = "0.4.34" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "0ceec5bc11778974d1bcb055b18002eba7f4b3518b6a0081b3af5f21666da9ad" +checksum = "f9f8bd3e56ce4dfc153cf470fffbfa98c7620958b312ca5c3a4b8d5181fd13c6" [[package]] name = "plain" @@ -156,9 +156,9 @@ checksum = "cab834c73d247e67f4fae452806d17d3c7501756d98c8808d7c9c7aa7d18f973" [[package]] name = "rustix" -version = "1.1.4" +version = "1.1.5" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "b6fe4565b9518b83ef4f91bb47ce29620ca828bd32cb7e408f0062e9930ba190" +checksum = "891efababe418670775f199f0d233d84843c227a0949a883ce15b37c78d6629d" dependencies = [ "bitflags", "errno", @@ -178,13 +178,13 @@ dependencies = [ [[package]] name = "scroll_derive" -version = "0.13.1" +version = "0.13.2" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "ed76efe62313ab6610570951494bdaa81568026e0318eaa55f167de70eeea67d" +checksum = "e1a36a382ed65dbcc0ab47fd5e9a94112417ccd34560a392ef3b7b0f0ec39148" dependencies = [ "proc-macro2", "quote", - "syn", + "syn 3.0.6", ] [[package]] @@ -198,6 +198,17 @@ dependencies = [ "unicode-ident", ] +[[package]] +name = "syn" +version = "3.0.6" +source = "registry+https://github.com/rust-lang/crates.io-index" +checksum = "8593e8e72159ed2257d083c7a454a85cbf854f37a0966d8d483aff8c8a3ebcee" +dependencies = [ + "proc-macro2", + "quote", + "unicode-ident", +] + [[package]] name = "terminal_size" version = "0.4.4" @@ -210,9 +221,9 @@ dependencies = [ [[package]] name = "unicode-ident" -version = "1.0.24" +version = "1.0.26" source = "registry+https://github.com/rust-lang/crates.io-index" -checksum = "e6e4313cd5fcd3dad5cafa179702e2b244f760991f45397d14d4ebf38247da75" +checksum = "d245f478577f809a851594d02313b640fb437e0bb33866753cff937863096954" [[package]] name = "windows-link" diff --git a/Cargo.toml b/Cargo.toml index b23f4de..6ae519a 100644 --- a/Cargo.toml +++ b/Cargo.toml @@ -1,6 +1,6 @@ [package] name = "gobin" -version = "0.5.0" +version = "0.5.1" edition = "2024" rust-version = "1.88" description = "Static analysis library for Go compiled binaries - identification and metadata extraction" @@ -28,7 +28,7 @@ indexing_slicing = "deny" [dependencies] # We only parse single ELF / Mach-O / PE executables (the three formats # dispatched in `formats.rs`), so drop goblin's default `te` (UEFI Terse -# Executable) feature — not a Go output format. (`archive` cannot be +# Executable) feature - not a Go output format. (`archive` cannot be # dropped: goblin's `mach32`/`mach64` features hard-require it.) PE # resource/import/cert/TLS *parsing* is disabled separately at the call # site via `ParseOptions`. `endian_fd` + `std` stay for sliced reads. diff --git a/benches/extract.rs b/benches/extract.rs index 4ae95fe..a9ee7da 100644 --- a/benches/extract.rs +++ b/benches/extract.rs @@ -72,7 +72,7 @@ fn path_of(name: &str) -> &'static str { /// Read a fixture once per process and hand out a shared borrow. /// /// Benchmarks measure gobin, not the filesystem, so the read must not land -/// inside the timed region — and it must not be re-counted by the allocation +/// inside the timed region - and it must not be re-counted by the allocation /// profiler on every iteration either. fn fixture(path: &str) -> &'static [u8] { static CACHE: OnceLock>> = OnceLock::new(); @@ -98,7 +98,7 @@ fn parse(bencher: Bencher, name: &str) { bencher.bench(|| black_box(GoBinary::parse(black_box(data))).is_some()); } -/// Function enumeration with names and source files resolved — the most +/// Function enumeration with names and source files resolved - the most /// commonly consumed surface. #[divan::bench(args = FIXTURES)] fn functions(bencher: Bencher, name: &str) { @@ -123,7 +123,7 @@ fn all_types(bencher: Bencher, name: &str) { bencher.bench(|| black_box(bin.all_types().len())); } -/// Go string-literal recovery — a full pointer-aligned sweep of the image. +/// Go string-literal recovery - a full pointer-aligned sweep of the image. #[divan::bench(args = FIXTURES)] fn strings(bencher: Bencher, name: &str) { let data = fixture(path_of(name)); @@ -139,7 +139,7 @@ fn itabs(bencher: Bencher, name: &str) { bencher.bench(|| black_box(bin.itab_pairs().count())); } -/// Inline-tree decoding over every function — the heaviest pclntab surface. +/// Inline-tree decoding over every function - the heaviest pclntab surface. #[divan::bench(args = FIXTURES)] fn inline_trees(bencher: Bencher, name: &str) { let data = fixture(path_of(name)); @@ -194,7 +194,7 @@ fn init_order(bencher: Bencher, name: &str) { bencher.bench(|| black_box(bin.init_order().len())); } -/// `//go:embed` payload recovery — a symbol-independent structural search. +/// `//go:embed` payload recovery - a symbol-independent structural search. #[divan::bench(args = FIXTURES)] fn embeds(bencher: Bencher, name: &str) { let data = fixture(path_of(name)); @@ -202,7 +202,7 @@ fn embeds(bencher: Bencher, name: &str) { bencher.bench(|| black_box(bin.embedded_assets().len())); } -/// A full metadata sweep — the shape of an actual triage run. +/// A full metadata sweep - the shape of an actual triage run. #[divan::bench(args = FIXTURES)] fn full_sweep(bencher: Bencher, name: &str) { let data = fixture(path_of(name)); diff --git a/src/detection.rs b/src/detection.rs index 06b6449..bab68d4 100644 --- a/src/detection.rs +++ b/src/detection.rs @@ -38,11 +38,11 @@ use crate::structures::PclntabVersion; /// Consumers persist this enum into long-lived schemas (database columns, /// structured logs). The contract: /// -/// - **Variants** — append-only. New tiers appear as new variants; existing +/// - **Variants** - append-only. New tiers appear as new variants; existing /// variants are never renamed or removed. -/// - **`Display` strings** (and [`Self::as_str`]) — fixed forever once +/// - **`Display` strings** (and [`Self::as_str`]) - fixed forever once /// shipped. Treat them as serialization keys. -/// - **`Debug` strings** — *not* a stability surface. Use `Display` / +/// - **`Debug` strings** - *not* a stability surface. Use `Display` / /// `as_str` for anything that lands in a database column. #[derive(Debug, Clone, Copy, PartialEq, Eq, PartialOrd, Ord)] pub enum Confidence { @@ -132,7 +132,7 @@ pub enum ConfidenceSignal { /// ELF `.note.go.buildid` (or `Go\0\0` note marker) was present. BuildidNotePresent, /// A Go type-metadata section was present. These names are unique to the - /// Go linker, so any of them is structural proof on its own — useful on + /// Go linker, so any of them is structural proof on its own - useful on /// stripped binaries where `.gopclntab` was renamed away. TypeSectionPresent { /// Which section matched, in its ELF spelling (`".typelink"`, diff --git a/src/formats.rs b/src/formats.rs index 7c1fce0..950ac46 100644 --- a/src/formats.rs +++ b/src/formats.rs @@ -51,7 +51,7 @@ use crate::{ }; /// Wasm section id of the Code section, which holds the module's function -/// bodies — the wasm equivalent of `.text`. +/// bodies - the wasm equivalent of `.text`. const WASM_CODE_SECTION_ID: u8 = 10; /// Wasm section id of the Data section, which holds every initialized byte of @@ -71,7 +71,7 @@ pub enum BinaryFormat { /// PE (Portable Executable) -- Windows. /// Magic: `MZ` (`4d 5a`) DOS header. Pe, - /// WebAssembly module (`.wasm`) — produced by `GOOS=js GOARCH=wasm` (or + /// WebAssembly module (`.wasm`) - produced by `GOOS=js GOARCH=wasm` (or /// `wasip1` since Go 1.21). /// Magic: `\0asm` + version `01 00 00 00`. /// @@ -80,7 +80,7 @@ pub enum BinaryFormat { /// translations make wasm look the same as the other formats to /// downstream parsers: /// - /// - `image_base` is `0` — linear-memory addresses are absolute. + /// - `image_base` is `0` - linear-memory addresses are absolute. /// - The Data section's many segments are reassembled into one /// contiguous linear-memory image (with zero-fill gaps), and that /// image becomes the address space [`BinaryContext::va_to_file`] @@ -89,11 +89,11 @@ pub enum BinaryFormat { /// in linear memory; the image presents them contiguously. /// /// The Go linker does not emit format-specific sections (`.gopclntab` - /// etc.) for wasm — only three custom sections: `go:buildid`, + /// etc.) for wasm - only three custom sections: `go:buildid`, /// `producers`, `name`. pclntab and buildinfo bytes live inside the /// Data-section linear-memory payload alongside the rest of the /// runtime's static data. `text_va` and `etext_va` here are runtime - /// "PC" boundaries — not byte offsets into the Code section — since + /// "PC" boundaries - not byte offsets into the Code section - since /// wasm encodes a Go PC as `(function_index << 16) | bytecode_offset`. Wasm, /// Unrecognized format. Magic-byte scanning can still find Go structures. @@ -121,7 +121,7 @@ pub struct GoSections { pub go_module: Option, /// File byte range of the typelink section (ELF / Mach-O, Go ≤ 1.26). /// Removed by Go 1.27, which replaced the typelink array with a walk over - /// the type-descriptor region — see [`Self::go_type`]. + /// the type-descriptor region - see [`Self::go_type`]. pub typelink: Option, /// File byte range of the itablink section (ELF / Mach-O, Go ≤ 1.26). /// Removed by Go 1.27, which stores itabs inline in the types region. @@ -133,7 +133,7 @@ pub struct GoSections { /// (which keeps everything in `.rdata`) and on wasm. pub go_type: Option, /// File byte range of the `go:funcdesc` section (`.go.func` / `__go_func`, - /// Go 1.27+) — the `·f` function-value descriptors that used to sit in + /// Go 1.27+) - the `·f` function-value descriptors that used to sit in /// `.rodata`. Recorded as a Go 1.27 structural marker. pub go_func: Option, /// File byte range of the FIPS-140 info section (`.go.fipsinfo` / @@ -146,7 +146,7 @@ pub struct GoSections { /// is otherwise a whole-image sweep. pub noptrdata: Option, /// File byte range of the initialized data section (`.data` / `__data`), - /// or — for wasm, which has no named sections — of the Data section + /// or - for wasm, which has no named sections - of the Data section /// payload. PE merges every Go data symbol into `.data`, so it is the PE /// equivalent of `noptrdata` for moduledata discovery. /// @@ -180,7 +180,7 @@ pub struct SectionRange { /// /// Parses the executable format **once** (via `goblin`) during construction and /// provides zero-copy section slicing, VA-to-file-offset translation, and ELF -/// note segment access. This is the low-level entry point — all Go metadata +/// note segment access. This is the low-level entry point - all Go metadata /// parsers (pclntab, buildinfo, types, etc.) receive a `&BinaryContext` rather /// than re-parsing the binary independently. /// @@ -203,7 +203,7 @@ pub struct BinaryContext<'a> { elf_note_segments: Vec<(usize, usize)>, /// Image base virtual address. /// - /// For PE binaries this is the `OptionalHeader.ImageBase` field — RVAs in + /// For PE binaries this is the `OptionalHeader.ImageBase` field - RVAs in /// the PE address space are relative to it. For ELF and Mach-O the field /// is `0`, since their addresses are already absolute VAs and "RVA" /// effectively coincides with VA. `Unknown` formats also report `0`. @@ -219,7 +219,7 @@ pub struct BinaryContext<'a> { /// /// Owned by [`BinaryContext`] (no leak). Borrows handed out via /// [`Self::structure_search_data`] are tied to `&self` and live as long - /// as the context does — all parsers that build structures borrowing + /// as the context does - all parsers that build structures borrowing /// from the LM image (`ParsedPclntab`, `GoType`, etc.) borrow with the /// same `&self` lifetime. wasm_lm: Option>, @@ -235,7 +235,7 @@ impl<'a> BinaryContext<'a> { /// Parse a binary, extracting format info, Go sections, VA mappings, and ELF notes /// in a single `goblin` pass. /// - /// Always succeeds — returns a context with empty sections/segments if `goblin` + /// Always succeeds - returns a context with empty sections/segments if `goblin` /// cannot parse the data. pub fn new(data: &'a [u8]) -> Self { let format = detect_format(data); @@ -259,7 +259,7 @@ impl<'a> BinaryContext<'a> { let mut elf_note_segments = Vec::new(); let mut image_base: u64 = 0; // Mach-O segment `(vmaddr, fileoff)` in load-command order, and the - // `LC_DYLD_CHAINED_FIXUPS` data offset — both needed to rebase chained + // `LC_DYLD_CHAINED_FIXUPS` data offset - both needed to rebase chained // fixups in externally-linked objects (plugins / CGO / c-shared). let mut macho_segments: Vec<(u64, u64)> = Vec::new(); let mut chained_fixups_off: Option = None; @@ -306,7 +306,7 @@ impl<'a> BinaryContext<'a> { } } BinaryFormat::MachO => { - // `MachO::parse(.., 0)` handles a thin (non-fat) Mach-O — the + // `MachO::parse(.., 0)` handles a thin (non-fat) Mach-O - the // only Mach variant the original `Mach::Binary` arm processed. if let Ok(macho) = MachO::parse(data, 0) { // Locate the chained-fixups load command, if present. @@ -316,7 +316,7 @@ impl<'a> BinaryContext<'a> { } } for seg in &macho.segments { - // Per-segment (vmaddr, fileoff) in load order — the + // Per-segment (vmaddr, fileoff) in load order - the // fixup chains index segments by this order. macho_segments.push((seg.vmaddr, seg.fileoff)); // VA mapping @@ -343,7 +343,7 @@ impl<'a> BinaryContext<'a> { BinaryFormat::Pe => { // We read only the optional header's image base and the // section table. Disable goblin's resource, import, - // certificate, and TLS parsing — Go binaries can carry large + // certificate, and TLS parsing - Go binaries can carry large // import/resource tables we never touch, and skipping them is // a straight efficiency win with no effect on what we use. let opts = ParseOptions::default() @@ -394,7 +394,7 @@ impl<'a> BinaryContext<'a> { for sec in walk_wasm_sections(data) { if sec.id == 0 && sec.name == Some("go:buildid") { // Presence is a hard signal that the Go linker produced - // this binary. Reuse the existing flag — buildid extract + // this binary. Reuse the existing flag - buildid extract // falls through to the raw-marker scan and finds the // payload bytes inside the section. sections.has_go_buildid_note = true; @@ -402,7 +402,7 @@ impl<'a> BinaryContext<'a> { if sec.id == WASM_CODE_SECTION_ID && sec.payload_size > 0 { // The Code section is wasm's executable region. Recorded // under `text_section` so the structural searches skip it - // exactly as they skip `.text` elsewhere — it is the + // exactly as they skip `.text` elsewhere - it is the // largest part of a Go wasm module and can hold none of // the data structures they look for. sections.text_section = Some(SectionRange { @@ -431,7 +431,7 @@ impl<'a> BinaryContext<'a> { // Register the linear-memory image as a single VA mapping. // This makes `va_to_file(va) == va`, so every code path that // does `structure_search_data()[va_to_file(va)..]` reads - // contiguous bytes across wasm data-segment boundaries — + // contiguous bytes across wasm data-segment boundaries - // gaps are zero-fill in the reconstructed image. let lm_len = image.len() as u64; segments.push((0u64, 0u64, lm_len)); @@ -466,7 +466,7 @@ impl<'a> BinaryContext<'a> { } } - /// Byte ranges of [`Self::data`] — i.e. **file** offsets — that can hold + /// Byte ranges of [`Self::data`] - i.e. **file** offsets - that can hold /// Go data structures, in order and non-overlapping. /// /// Callers that search through [`Self::structure_search_data`] want @@ -474,7 +474,7 @@ impl<'a> BinaryContext<'a> { /// offsets. /// /// Several extraction surfaces have no symbol to look up and must search - /// memory structurally — the build-info blob, `embed.FS` file arrays, the + /// memory structurally - the build-info blob, `embed.FS` file arrays, the /// moduledata. All of them are *data*, so the two largest regions of a Go /// binary can be excluded outright: the executable section (`.text`, /// `__text`, or a wasm Code section) and the pclntab, which is a @@ -508,13 +508,13 @@ impl<'a> BinaryContext<'a> { } /// The same idea as [`Self::data_regions`], but as ranges into - /// [`Self::structure_search_data`] — the view every structural parser + /// [`Self::structure_search_data`] - the view every structural parser /// actually reads through. /// /// For ELF, Mach-O and PE that view is the file (or, for a chained-fixup /// Mach-O, a rebased copy with identical layout), so the file ranges carry /// over unchanged. For wasm it is the reconstructed linear-memory image, - /// whose offsets are linear-memory addresses unrelated to file positions — + /// whose offsets are linear-memory addresses unrelated to file positions - /// and which is *entirely* initialized data, so there is nothing to /// exclude. pub fn search_regions(&self) -> Vec<(usize, usize)> { @@ -585,7 +585,7 @@ impl<'a> BinaryContext<'a> { /// /// For wasm this is the reconstructed linear-memory image (covering all /// data segments laid out at their target offsets, with zero-fill in - /// between). For every other format this is the file bytes — wasm is the + /// between). For every other format this is the file bytes - wasm is the /// only one where Go runtime structures span multiple disjoint regions. /// /// Offsets returned by [`Self::va_to_file`] index into this slice for @@ -684,7 +684,7 @@ pub fn detect_format(data: &[u8]) -> BinaryFormat { /// rename every read-only-relocatable Go section with a `.data.rel.ro` prefix /// (`cmd/link/internal/ld/data.go`, `genrelrosecname`), so `.typelink` becomes /// `.data.rel.ro.typelink` and `.go.type` becomes `.data.rel.ro.go.type`. The -/// prefix is stripped before matching — otherwise a PIE binary looks like it +/// prefix is stripped before matching - otherwise a PIE binary looks like it /// has no typelink section at all, which the moduledata layout arbitration /// reads as a Go 1.27 signal. fn classify_section(name: &str, range: Option, result: &mut GoSections) { diff --git a/src/lib.rs b/src/lib.rs index c674c6a..aa7b802 100644 --- a/src/lib.rs +++ b/src/lib.rs @@ -44,7 +44,7 @@ //! Wasm support reconstructs a single linear-memory image from the wasm //! Data section's individual segments so runtime structures (pclntab, //! moduledata, type descriptors) that span multiple disjoint segments can -//! be addressed by their linear-memory VA — see +//! be addressed by their linear-memory VA - see //! [`structures::wasm`] and the [`BinaryFormat::Wasm`] //! variant rustdoc for details. //! @@ -120,7 +120,7 @@ pub struct GoBinary<'a> { build_info: Option>, /// Cached pclntab scalars. The borrowing [`ParsedPclntab`] view is /// reconstructed on demand via [`Self::pclntab`] so it can borrow from - /// `&self.ctx` — important for wasm, where the address space pclntab + /// `&self.ctx` - important for wasm, where the address space pclntab /// lives in is the linear-memory image owned by the context (and is /// therefore not borrowable with the input lifetime `'a`). pclntab_meta: Option, @@ -146,7 +146,7 @@ impl<'a> GoBinary<'a> { /// # Working with mmap-ed input /// /// `parse` borrows the input for the lifetime of the returned [`GoBinary`]. - /// Any byte slice works — including one backed by `memmap2::Mmap` — + /// Any byte slice works - including one backed by `memmap2::Mmap` - /// regardless of source. There is no separate mmap-specific entry point; /// just pass `&mmap[..]`. /// @@ -161,8 +161,8 @@ impl<'a> GoBinary<'a> { /// containing the [`ConfidenceReport`] gathered during detection. /// /// Detection signals are also retained on success, accessible via - /// [`Self::report`] — useful for surfacing analyst-facing diagnostics - /// (e.g. "Go binary, but pclntab missing — likely heavily patched"). + /// [`Self::report`] - useful for surfacing analyst-facing diagnostics + /// (e.g. "Go binary, but pclntab missing - likely heavily patched"). pub fn try_parse(data: &'a [u8]) -> Result { let ctx = BinaryContext::new(data); let mut report = ConfidenceReport::empty(); @@ -289,11 +289,11 @@ impl<'a> GoBinary<'a> { self.report.tier } - /// Structured detection report — the confidence tier plus every signal + /// Structured detection report - the confidence tier plus every signal /// observed during parse. /// /// Useful for analyst-facing diagnostics ("Go binary, but pclntab is - /// missing — likely heavily patched") and for surfacing details in bug + /// missing - likely heavily patched") and for surfacing details in bug /// reports. pub fn report(&self) -> &ConfidenceReport { &self.report @@ -342,7 +342,7 @@ impl<'a> GoBinary<'a> { /// /// For bulk per-function processing where you also need decoded pcsp / /// pcln / pcfile tables, use [`crate::metadata::for_each_function`] - /// instead — it amortizes table-decode buffers across the whole walk. + /// instead - it amortizes table-decode buffers across the whole walk. pub fn functions(&self) -> FunctionIter<'_> { FunctionIter::new(self.pclntab()) } @@ -381,7 +381,7 @@ impl<'a> GoBinary<'a> { /// The runtime links one `moduledata` per loaded module (the main binary /// plus any plugins / shared libraries) via the `next` field. **That field /// is populated at load time**, so in a static on-disk binary it is almost - /// always nil and this returns a single entry — the chain is followed + /// always nil and this returns a single entry - the chain is followed /// defensively (with a cycle guard) for the rare multi-module image and so /// plugin/shared objects analyzed on their own parse correctly. Empty if /// the binary has no locatable moduledata. @@ -419,7 +419,7 @@ impl<'a> GoBinary<'a> { Moduledata::parse(bytes, meta.ptr_size, hints) } - /// Virtual address of `runtime.text` — the first byte of Go-emitted code. + /// Virtual address of `runtime.text` - the first byte of Go-emitted code. /// /// `entry_off` on each [`FuncData`] is measured relative to this address. /// In most cases callers should reach for [`Self::entry_va`] / @@ -433,7 +433,7 @@ impl<'a> GoBinary<'a> { /// tried in decreasing order of authority: /// /// 1. `moduledata.text`. - /// 2. The pclntab's own record of `runtime.text` — the pcHeader + /// 2. The pclntab's own record of `runtime.text` - the pcHeader /// `textStart` field on Go 1.18-1.25, or the lowest absolute function /// PC on Go 1.2-1.17 (both surface as `header_text_start` / /// `text_start`). @@ -443,7 +443,7 @@ impl<'a> GoBinary<'a> { /// `runtime.text` at the start of `.text` / `__text` on every format, /// so the section header supplies it. /// - /// Returns `None` rather than `0` when nothing is available — a bogus + /// Returns `None` rather than `0` when nothing is available - a bogus /// `Some(0)` would silently turn every [`Self::entry_va`] into a raw /// `entry_off`. pub fn text_va(&self) -> Option { @@ -475,7 +475,7 @@ impl<'a> GoBinary<'a> { /// (binary lacks moduledata or VA mapping) or the addition overflows. /// /// For ELF and Mach-O this is the address a disassembler will use - /// directly. For PE this is still a true VA — pass it to + /// directly. For PE this is still a true VA - pass it to /// [`Self::entry_rva`] (or subtract [`BinaryContext::image_base`]) to get /// the RVA most PE-aware tools expect. pub fn entry_va(&self, func: &FuncData) -> Option { @@ -494,7 +494,7 @@ impl<'a> GoBinary<'a> { self.entry_va(func)?.checked_sub(self.ctx.image_base()) } - /// Virtual address of `runtime.etext` — one past the last byte of + /// Virtual address of `runtime.etext` - one past the last byte of /// Go-emitted code. /// /// `etext_va() - text_va()` gives the total size of all Go-emitted code, @@ -503,7 +503,7 @@ impl<'a> GoBinary<'a> { self.moduledata.as_ref().map(|m| m.etext) } - /// Whether this module contains the program's `main` — true for the main + /// Whether this module contains the program's `main` - true for the main /// executable, false for a plugin / shared library. /// /// Read from `moduledata.hasmain`. `None` if moduledata is unavailable. @@ -511,12 +511,12 @@ impl<'a> GoBinary<'a> { self.moduledata.as_ref().map(|m| m.has_main) } - /// Pointer map of the initialized data segment (`[data, edata)`) — which + /// Pointer map of the initialized data segment (`[data, edata)`) - which /// pointer-sized words hold pointers, decoded from the `gcdata` GC program. /// /// This is the precise location of every pointer in global initialized /// memory (function pointers, `itab`/interface pointers, string/slice - /// headers, global `*T` variables) — recoverable without disassembly. + /// headers, global `*T` variables) - recoverable without disassembly. /// `None` if moduledata / the GC data is unavailable. pub fn data_pointer_map(&self) -> Option { let md = self.moduledata.as_ref()?; @@ -565,7 +565,7 @@ impl<'a> GoBinary<'a> { } /// Whether the binary was built with coverage instrumentation - /// (`-cover`) — detected via a non-empty `moduledata` coverage-counter + /// (`-cover`) - detected via a non-empty `moduledata` coverage-counter /// region. Always `false` for Go < 1.20 (the region did not exist). pub fn is_coverage_build(&self) -> bool { self.moduledata @@ -711,7 +711,7 @@ impl<'a> GoBinary<'a> { /// Which Go compiler toolchain produced this binary. /// /// Detection order: - /// 1. `-compiler` build setting (`gc`, `gccgo`, etc.) — authoritative. + /// 1. `-compiler` build setting (`gc`, `gccgo`, etc.) - authoritative. /// 2. `tinygo` substring in the Go version string. /// 3. Presence of pclntab → `gc` (TinyGo and gccgo do not produce it). /// 4. Otherwise [`Compiler::Unknown`]. @@ -795,7 +795,7 @@ impl<'a> GoBinary<'a> { /// commit hash is not stamped into the version string. /// /// For CVE matching against the Go toolchain itself, the commit hash is - /// more precise than the marketing version — released versions only narrow + /// more precise than the marketing version - released versions only narrow /// to a tag. pub fn runtime_commit(&self) -> Option<&str> { let v = self.go_version?; @@ -867,7 +867,7 @@ impl<'a> GoBinary<'a> { /// and backing bytes (borrowed from read-only data). Works on stripped /// binaries. Returns an empty `Vec` when the binary embeds nothing. /// - /// Only the multi-file `embed.FS` form is recovered — the single-file + /// Only the multi-file `embed.FS` form is recovered - the single-file /// `//go:embed` string/`[]byte` forms compile to plain variables with no /// recognizable anchor and are not surfaced. This is the path Visus uses /// to recurse embedded dropper payloads out of Go binaries. @@ -887,7 +887,7 @@ impl<'a> GoBinary<'a> { /// binary predates Go 1.24, lacks moduledata, or carries no init tasks. /// /// Init order reveals which packages run setup code at startup and in what - /// sequence — a useful lens on staging / persistence behaviour. + /// sequence - a useful lens on staging / persistence behaviour. pub fn init_order(&self) -> Vec> { let md = match self.moduledata.as_ref() { Some(m) => m, @@ -903,7 +903,7 @@ impl<'a> GoBinary<'a> { // // A binary has thousands of functions and a couple of dozen init // tasks, so the pass collects only the entry offsets the tasks - // actually reference — indexing every function into a map costs more + // actually reference - indexing every function into a map costs more // memory than the whole rest of this call and throws almost all of it // away. It also walks `func_entries` rather than `functions()`, which // would additionally decode a source file, line range and frame size @@ -971,14 +971,14 @@ impl<'a> GoBinary<'a> { /// `types+itaboffset` (Go 1.27+ / V5). Returns an empty iterator when no /// source is available (heavily stripped binaries). /// - /// Useful for "what implements `io.Reader` in this binary?" queries — + /// Useful for "what implements `io.Reader` in this binary?" queries - /// pair with [`Self::types`] to resolve each VA back to a named type. pub fn itab_pairs(&self) -> itab::ItabIter<'_> { let ptr_size = self.pclntab_meta.map(|m| m.ptr_size).unwrap_or(0); itab::extract_iter(&self.ctx, ptr_size, self.moduledata.as_ref()) } - /// The `Fun[]` method-pointer array of an itab — the concrete-type methods + /// The `Fun[]` method-pointer array of an itab - the concrete-type methods /// bound to each interface method, in interface order (a `0` entry means /// the method is unbound). Pair with [`Self::itab_pairs`]. pub fn itab_methods(&self, pair: &itab::ItabPair) -> Vec { @@ -989,13 +989,13 @@ impl<'a> GoBinary<'a> { /// Whether the binary's pclntab references any cgo-related runtime /// functions (`runtime.cgocall`, `runtime.cgocallback`, etc.). /// - /// This is a binary-level "did this binary use cgo at all?" signal — a + /// This is a binary-level "did this binary use cgo at all?" signal - a /// strong indicator the program may execute native code from C (DLLs, /// syscalls, exploits). Per-call-site enumeration would require /// disassembly support, which the crate does not have today. pub fn has_cgo(&self) -> bool { // Short-circuits on the first matching function via the streaming - // iterator — does not materialize the whole function list. + // iterator - does not materialize the whole function list. self.functions().any(|f| is_cgo_runtime_fn(f.name)) } @@ -1050,7 +1050,7 @@ impl<'a> GoBinary<'a> { /// (pointer/slice/array/chan elements, map key/value, struct field types, /// func parameter/result types, method signatures, and each type's /// pointer-to-this), parsing each descriptor independently by virtual - /// address. This reaches types absent from `typelink` — e.g. a struct type + /// address. This reaches types absent from `typelink` - e.g. a struct type /// used only as a pointer's element, together with its field tags. Capped /// to bound pathological graphs. pub fn all_types(&self) -> Vec> { @@ -1100,12 +1100,12 @@ impl<'a> GoBinary<'a> { /// on Go binaries. /// /// Yields zero items when the binary lacks VA mapping. Length filter: - /// 2..=4096 bytes. UTF-8 is **not** required — malware frequently stashes + /// 2..=4096 bytes. UTF-8 is **not** required - malware frequently stashes /// non-UTF-8 payloads in length-prefixed rodata entries; use /// [`gostrings::GoString::as_bytes`] for raw bytes, /// [`gostrings::GoString::try_as_str`] / [`gostrings::GoString::as_str`] /// for text. Pointers into the text segment (`[moduledata.text, - /// moduledata.etext)`) are excluded. **Duplicates are not filtered** — + /// moduledata.etext)`) are excluded. **Duplicates are not filtered** - /// a string referenced from N positions yields N times. Collect into a /// `HashSet` if you want unique results. pub fn strings(&self) -> gostrings::GoStringIter<'_> { @@ -1128,7 +1128,7 @@ impl<'a> GoBinary<'a> { /// invoke the full analyzer. /// /// False negatives are possible (heavily patched binaries where every marker -/// has been wiped). False positives are unlikely — these magic byte sequences +/// has been wiped). False positives are unlikely - these magic byte sequences /// don't naturally appear in non-Go binaries. pub fn detect(data: &[u8]) -> bool { if find_bytes(data, b"\xff Go buildinf:").is_some() { @@ -1281,7 +1281,7 @@ fn parse_go_minor_version(version: &str) -> Option { /// Locate and parse the moduledata for accessor-only use (text/etext/types /// region addresses). /// -/// Delegates to [`ModuledataLocator`], which owns every discovery strategy — +/// Delegates to [`ModuledataLocator`], which owns every discovery strategy - /// the `.go.module` section on Go 1.26+ ELF / Mach-O, and the `pcHeader`- /// pointer scan everywhere else. Returns `None` if the binary lacks VA /// mappings or the moduledata cannot be located; callers degrade gracefully diff --git a/src/metadata.rs b/src/metadata.rs index a4ce520..1badb59 100644 --- a/src/metadata.rs +++ b/src/metadata.rs @@ -41,7 +41,7 @@ use crate::{ /// /// Covers all three spellings the runtime has used: `runtime` itself, /// `runtime/` (e.g. `runtime/cgo`, and the pre-1.24 `runtime/internal/*` -/// tree), and `internal/runtime/` — the home the runtime's internal +/// tree), and `internal/runtime/` - the home the runtime's internal /// packages (`internal/runtime/atomic`, `internal/runtime/maps`, /// `internal/runtime/sys`, …) moved to in Go 1.24 and where they still live in /// 1.27. Without the third form, half the runtime of a modern binary @@ -83,10 +83,10 @@ pub fn is_stdlib_path(pkg: &str) -> bool { pub enum Compiler { /// The standard `gc` compiler (the default Go toolchain). Gc, - /// TinyGo — produces small embedded/wasm binaries with a different runtime. + /// TinyGo - produces small embedded/wasm binaries with a different runtime. /// TinyGo binaries do not carry a stdlib pclntab. TinyGo, - /// Gccgo — GCC's Go front-end. Produces no pclntab. + /// Gccgo - GCC's Go front-end. Produces no pclntab. Gccgo, /// Could not determine (no `-compiler` setting and no distinguishing /// markers were found). @@ -111,7 +111,7 @@ pub struct DepEntry<'a> { /// Replacement target for a [`DepEntry`]. /// -/// Records what the original module was substituted with — either a forked +/// Records what the original module was substituted with - either a forked /// module (`=> github.com/forked/foo v1.2.3 h1:xyz=`) or a local path /// (`=> ./local/foo`, in which case `version` is typically `None`). #[derive(Debug, Clone, PartialEq, Eq)] @@ -134,11 +134,11 @@ pub struct DepReplacement<'a> { /// Consumers persist this enum into long-lived schemas (database columns, /// structured logs). The contract: /// -/// - **Variants** — append-only. New obfuscators appear as new variants; +/// - **Variants** - append-only. New obfuscators appear as new variants; /// existing variants are never renamed or removed. -/// - **`Display` strings** (and [`Self::kind_str`]) — fixed forever once +/// - **`Display` strings** (and [`Self::kind_str`]) - fixed forever once /// shipped. Treat them as serialization keys. -/// - **`Debug` strings** — *not* a stability surface. Use `Display` / +/// - **`Debug` strings** - *not* a stability surface. Use `Display` / /// `kind_str` for anything that lands in a database column. #[derive(Debug, Clone, PartialEq, Eq)] pub enum ObfuscationKind { @@ -162,7 +162,7 @@ impl ObfuscationKind { /// Stable lowercase identifier for this verdict /// (`"none"` / `"garble"` / `"other"`). /// - /// Drops the inner `confidence` and `reason` payloads — read those off + /// Drops the inner `confidence` and `reason` payloads - read those off /// the matched variant directly. See the `# Stability` section on /// [`ObfuscationKind`] for the durability contract. pub fn kind_str(&self) -> &'static str { @@ -190,7 +190,7 @@ impl std::fmt::Display for ObfuscationKind { /// setting, recorded in build info. The `__go_fipsinfo` section's integrity /// sum (the `go:fipsinfo` symbol, `struct { Magic [16]byte; Sum [32]byte }`) /// is present in *every* crypto-linked binary, so it alone does **not** -/// indicate FIPS mode — it's exposed here only as a content identifier for +/// indicate FIPS mode - it's exposed here only as a content identifier for /// the embedded crypto module. /// /// [`crate::GoBinary::fips_info`] returns this only when `GOFIPS140` is set; @@ -204,7 +204,7 @@ pub struct FipsInfo<'a> { /// The `GOFIPS140` build-setting value (e.g. `"v1.0.0"` or /// `"v1.0.0-c2097c7c"`). pub version: &'a str, - /// Whether `fips140=on` appears in `DefaultGODEBUG` — i.e. FIPS + /// Whether `fips140=on` appears in `DefaultGODEBUG` - i.e. FIPS /// enforcement is active by default at runtime (vs. `fips140=only`/opt-in). pub enforced_by_default: bool, /// The 32-byte integrity sum from `__go_fipsinfo`, when the section is @@ -473,12 +473,12 @@ fn package_boundary(name: &str) -> Option { /// Consumers persist this enum into long-lived schemas (database columns, /// structured logs). The contract: /// -/// - **Variants** — append-only. New build modes appear as new variants; +/// - **Variants** - append-only. New build modes appear as new variants; /// existing variants are never renamed or removed. -/// - **`Display` strings** (and [`Self::as_str`]) — fixed forever once +/// - **`Display` strings** (and [`Self::as_str`]) - fixed forever once /// shipped. They round-trip through [`Self::parse`]; treat them as /// serialization keys. -/// - **`Debug` strings** — *not* a stability surface. Use `Display` / +/// - **`Debug` strings** - *not* a stability surface. Use `Display` / /// `as_str` for anything that lands in a database column. #[derive(Debug, Clone, PartialEq, Eq)] pub enum BuildMode { @@ -640,7 +640,7 @@ impl<'a> BuildInfo<'a> { /// any active `replace` directive (`=> path[ version][ h1:sum]`). /// /// This surfaces the supply-chain detail that [`Self::dependencies`] - /// collapses — e.g. answering "is this binary using the official + /// collapses - e.g. answering "is this binary using the official /// `golang.org/x/crypto` or a forked/replaced one?". pub fn deps_full(&self) -> impl Iterator> + '_ { self.deps.iter() @@ -752,7 +752,7 @@ impl FunctionInfo<'_> { /// The short name (without package prefix). /// - /// Mirrors [`Self::package`] — both use the same boundary computation. + /// Mirrors [`Self::package`] - both use the same boundary computation. /// E.g. `"github.com/spf13/cobra.(*Command).Run"` -> `"(*Command).Run"`, /// `"gopkg.in/yaml.v3.Marshal"` -> `"Marshal"`. pub fn short_name(&self) -> &str { @@ -765,14 +765,14 @@ impl FunctionInfo<'_> { /// Whether this function is a method (has a receiver type). /// /// A method has the form `..` where - /// `` is either `(*?Type[generics?])` (parenthesized — pointer + /// `` is either `(*?Type[generics?])` (parenthesized - pointer /// or complex receiver) or a bare identifier with optional generic args. /// This parses the structure rather than using a substring heuristic, so /// it correctly identifies value-receiver methods like `time.Time.String` /// that the old `".("` heuristic missed. /// /// Closures (`pkg.parent.funcN`) and gowrap stubs (`pkg.parent.gowrapN`) - /// look structurally like methods but are excluded — see [`Self::is_closure`]. + /// look structurally like methods but are excluded - see [`Self::is_closure`]. pub fn is_method(&self) -> bool { if self.is_closure() { return false; @@ -854,7 +854,7 @@ impl FunctionInfo<'_> { /// Whether this function runs on the system stack (`FuncIDsystemstack` / /// `FuncIDsystemstack_switch`). /// - /// Inferred from [`Self::func_id`], not from `flag` — the runtime tracks + /// Inferred from [`Self::func_id`], not from `flag` - the runtime tracks /// systemstack by ID rather than a flag bit. pub fn is_systemstack(&self) -> bool { matches!(self.func_id, 98 | 99) @@ -865,7 +865,7 @@ impl FunctionInfo<'_> { /// The Go compiler emits closures with names of the form /// `parent.funcN` (and similarly `parent.gowrapN` for goroutine wrappers /// around method calls), where `N` is a positive integer. This checks - /// the structural suffix shape — not the substring `.func` — and excludes + /// the structural suffix shape - not the substring `.func` - and excludes /// hand-written assembly (which never produces closures and could /// otherwise share textual patterns). pub fn is_closure(&self) -> bool { @@ -917,7 +917,7 @@ impl FunctionInfo<'_> { /// Strongly-typed view of the `_func.flag` byte. /// -/// Source: `src/internal/abi/symtab.go` — three flag bits are defined as of +/// Source: `src/internal/abi/symtab.go` - three flag bits are defined as of /// Go 1.26: /// /// | Bit | Constant | Meaning | @@ -981,7 +981,7 @@ pub struct FunctionTables<'a> { /// a [`FunctionInfo`] and its decoded per-PC tables. /// /// Bulk equivalent of [`FunctionIter`] paired with per-function table -/// decoding — but using three reusable buffers shared across the whole walk +/// decoding - but using three reusable buffers shared across the whole walk /// instead of allocating fresh `Vec`s for each function. For binaries with /// tens of thousands of functions this avoids `O(nfunc)` allocations. /// @@ -1066,7 +1066,7 @@ where /// - `'a`: lifetime of the underlying binary bytes; yielded /// [`FunctionInfo`] structs borrow strings from there. /// -/// Skips functions whose `_func` struct fails to parse — adversarial pclntab +/// Skips functions whose `_func` struct fails to parse - adversarial pclntab /// data cannot panic the iteration. Yields zero items for binaries without a /// recoverable pclntab. pub struct FunctionIter<'a> { @@ -1364,7 +1364,7 @@ mod tests { assert!(make("main.main.func1").is_closure()); assert!(make("main.main.func1.func2").is_closure()); assert!(make("main.run.gowrap1").is_closure()); - // Just `.func` without a digit — not a closure + // Just `.func` without a digit - not a closure assert!(!make("pkg.Func").is_closure()); // A type literally named `Func` with a method assert!(!make("pkg.Func.Method").is_closure()); diff --git a/src/structures/abitype.rs b/src/structures/abitype.rs index dee9499..ff0fa2e 100644 --- a/src/structures/abitype.rs +++ b/src/structures/abitype.rs @@ -16,13 +16,13 @@ //! - `FieldAlign_` (u8) //! - `Kind_` (u8) //! - `Equal` (uintptr, equality function pointer) -//! - `GCData` (uintptr, GC pointer-mask bitmap — or a pointer to one; see +//! - `GCData` (uintptr, GC pointer-mask bitmap - or a pointer to one; see //! [`TFLAG_GC_MASK_ON_DEMAND`]) //! - `Str` (NameOff / i32) //! - `PtrToThis` (TypeOff / i32) //! //! Source: `src/internal/abi/type.go`. The struct itself has been stable since -//! Go 1.21; what changed in Go 1.27 is how `GCData` is populated — see +//! Go 1.21; what changed in Go 1.27 is how `GCData` is populated - see //! [`TFLAG_GC_MASK_ON_DEMAND`]. use crate::structures::util::{read_i32, read_u32, read_uintptr}; @@ -40,7 +40,7 @@ pub const TFLAG_NAMED: u8 = 0x04; /// memory (`TFlagRegularMemory`). pub const TFLAG_REGULAR_MEMORY: u8 = 0x08; -/// `TFlag` bit: `GCData` is **not** a pointer mask but a `**byte` — a slot the +/// `TFlag` bit: `GCData` is **not** a pointer mask but a `**byte` - a slot the /// runtime fills in with a lazily-built mask on first use /// (`runtime.getGCMaskOnDemand`). /// diff --git a/src/structures/buildinfo.rs b/src/structures/buildinfo.rs index ab2545e..4cbb0ed 100644 --- a/src/structures/buildinfo.rs +++ b/src/structures/buildinfo.rs @@ -142,9 +142,9 @@ pub fn extract<'a>(ctx: &BinaryContext<'a>) -> Option> { /// narrows to the regions that can hold it before falling back to the whole /// image: /// -/// 1. the dedicated `.go.buildinfo` / `__go_buildinfo` section — exact, and +/// 1. the dedicated `.go.buildinfo` / `__go_buildinfo` section - exact, and /// the usual case for ELF and Mach-O; -/// 2. the `.data` / `.noptrdata` sections — PE merges every Go data symbol +/// 2. the `.data` / `.noptrdata` sections - PE merges every Go data symbol /// into `.data`, which is a small fraction of a Go binary (tens of KB /// against megabytes of `.text` and `.rdata`); /// 3. the data regions of the image @@ -168,8 +168,8 @@ fn find_magic(ctx: &BinaryContext<'_>, data: &[u8]) -> Option { // Fall back to a sweep, but only over the regions that can hold data: the // blob is a `sym.SBUILDINFO` symbol, so the executable section and the // pclntab are excluded. That matters most for wasm, which names no - // sections at all and would otherwise sweep the whole module — twice, once - // per candidate list above — to prove a blob it never carries is absent. + // sections at all and would otherwise sweep the whole module - twice, once + // per candidate list above - to prove a blob it never carries is absent. for (from, to) in ctx.data_regions() { if let Some(region) = data.get(from..to) && let Some(pos) = find_aligned_magic(region) @@ -185,7 +185,7 @@ fn find_magic(ctx: &BinaryContext<'_>, data: &[u8]) -> Option { /// The Go linker aligns the symbol to [`BUILDINFO_ALIGN`] (a macOS /// requirement), so an aligned occurrence is the real one; an unaligned hit is /// still returned as a fallback in case a section offset shifted the -/// alignment. Both are answered in a **single** sweep — the previous +/// alignment. Both are answered in a **single** sweep - the previous /// aligned-then-unaligned pair walked the buffer twice, which on PE and wasm /// (neither of which has a `.go.buildinfo` section to narrow the search) meant /// scanning the whole image twice over. diff --git a/src/structures/embed.rs b/src/structures/embed.rs index 26bab9f..7d23a72 100644 --- a/src/structures/embed.rs +++ b/src/structures/embed.rs @@ -93,8 +93,8 @@ pub fn extract<'a>(ctx: &'a BinaryContext<'a>, ptr_size: u8) -> Vec( ctx: &'a BinaryContext<'a>, @@ -222,7 +222,7 @@ fn parse_file_array<'a>( /// /// Mirrors `embed.split`: strip a trailing `/`, then split at the last /// remaining `/` (a missing dir becomes `"."`). Borrowed from `name` rather -/// than owned — the key exists only to be compared against its predecessor, +/// than owned - the key exists only to be compared against its predecessor, /// and the blind scan evaluates it for every candidate entry it examines, so /// owning it allocated twice per rejected entry. fn embed_sort_key(name: &str) -> (&str, &str) { diff --git a/src/structures/fixups.rs b/src/structures/fixups.rs index 79a0d10..35b1cab 100644 --- a/src/structures/fixups.rs +++ b/src/structures/fixups.rs @@ -4,7 +4,7 @@ //! every data pointer not as a plain VA but as a link in a *fixup chain*: the //! 64-bit slot packs the real target into low bits plus chain-walk metadata //! (`next`, `bind`) in the high bits, to be applied by `dyld` at load time. -//! Until they are applied, reading such a slot as a pointer yields garbage — +//! Until they are applied, reading such a slot as a pointer yields garbage - //! so type / itab / moduledata pointer resolution fails on these objects. //! //! This module walks the fixup chains and produces a **rebased copy** of the @@ -18,9 +18,9 @@ use crate::structures::util::{read_u16, read_u32, read_uintptr}; -/// `DYLD_CHAINED_PTR_64` — 8-byte slots, `target` is an unslid vmaddr. +/// `DYLD_CHAINED_PTR_64` - 8-byte slots, `target` is an unslid vmaddr. const PTR_64: u16 = 2; -/// `DYLD_CHAINED_PTR_64_OFFSET` — 8-byte slots, `target` is an offset from the +/// `DYLD_CHAINED_PTR_64_OFFSET` - 8-byte slots, `target` is an offset from the /// image base. This is what the Go/clang toolchain emits. const PTR_64_OFFSET: u16 = 6; /// Sentinel page-start value meaning "no fixups on this page". @@ -37,7 +37,7 @@ fn read_u64(data: &[u8], off: usize) -> Option { /// that order). `image_base` is the lowest segment vmaddr. /// /// Returns `None` if there are no applicable (64-bit) fixups or the blob is -/// malformed — the caller then uses the original bytes unchanged. Never +/// malformed - the caller then uses the original bytes unchanged. Never /// panics; bounded against malformed offsets and chain cycles. pub fn rebase( data: &[u8], diff --git a/src/structures/gcprog.rs b/src/structures/gcprog.rs index 0de27dd..478b107 100644 --- a/src/structures/gcprog.rs +++ b/src/structures/gcprog.rs @@ -1,14 +1,14 @@ //! GC pointer-map (GC program) decoder. //! -//! `moduledata.gcdata` / `moduledata.gcbss` point at **GC programs** — a +//! `moduledata.gcdata` / `moduledata.gcbss` point at **GC programs** - a //! Lempel-Ziv-style bytecode the runtime expands (`runGCProg`) into a 1-bit //! per pointer-sized word bitmap over the `[data, edata)` / `[bss, ebss)` //! segments. Bit `i` set means the word at `segment_start + i*ptrSize` holds a //! pointer. //! //! Decoding this gives a precise map of **where pointers live in global -//! memory** — function pointers, interface/`itab` pointers, string/slice -//! headers, global `*T` variables — without any disassembly. +//! memory** - function pointers, interface/`itab` pointers, string/slice +//! headers, global `*T` variables - without any disassembly. //! //! ## Bytecode //! @@ -32,7 +32,7 @@ const HARD_WORD_CAP: usize = 64 * 1024 * 1024; /// `max_words` is the number of pointer-sized words the segment spans /// (`(edata - data) / ptrSize`). The returned vector has length `<= max_words`; /// element `i` is `true` when word `i` of the segment holds a pointer. Never -/// panics on malformed input — it stops early and returns what it decoded. +/// panics on malformed input - it stops early and returns what it decoded. pub fn run_gc_prog(prog: &[u8], max_words: usize) -> Vec { let cap = max_words.min(HARD_WORD_CAP); let mut out: Vec = Vec::new(); diff --git a/src/structures/gostring.rs b/src/structures/gostring.rs index f298182..8c772b3 100644 --- a/src/structures/gostring.rs +++ b/src/structures/gostring.rs @@ -1,4 +1,4 @@ -//! Go string header (`reflect.StringHeader`) layout — `(ptr, len)`, each +//! Go string header (`reflect.StringHeader`) layout - `(ptr, len)`, each //! `uintptr`-sized. Unlike slices, strings have no capacity field. //! //! This module provides the low-level *header* type used by other parsers. diff --git a/src/structures/inittask.rs b/src/structures/inittask.rs index cd39415..9057221 100644 --- a/src/structures/inittask.rs +++ b/src/structures/inittask.rs @@ -1,4 +1,4 @@ -//! `inittasks` decoder — recovers package initialization order. +//! `inittasks` decoder - recovers package initialization order. //! //! The linker builds `moduledata.inittasks` (`[]*initTask`, Go 1.24+): the //! ordered list of package-init work the runtime runs at startup. Each entry @@ -16,7 +16,7 @@ //! //! Each function pointer is the entry VA of an init function (`pkg.init`, //! `pkg.init.0`, …). Resolving those VAs back to names (done by the caller via -//! the pclntab function table) yields a readable startup order — useful for +//! the pclntab function table) yields a readable startup order - useful for //! understanding droppers' persistence / staging behaviour. use crate::{ diff --git a/src/structures/inline.rs b/src/structures/inline.rs index 9c82e42..7f661b1 100644 --- a/src/structures/inline.rs +++ b/src/structures/inline.rs @@ -21,9 +21,9 @@ //! //! ## Decoding chain //! -//! 1. Read `funcdata[FUNCDATA_InlTree]` (constant `3`) for the function — a +//! 1. Read `funcdata[FUNCDATA_InlTree]` (constant `3`) for the function - a //! `u32` offset added to `moduledata.gofunc` to get the inline-tree blob's VA. -//! 2. Decode `pcdata[PCDATA_InlTreeIndex]` (constant `2`) — yields +//! 2. Decode `pcdata[PCDATA_InlTreeIndex]` (constant `2`) - yields //! `(pc_offset, index)` pairs. `index < 0` means "not inlined here"; //! `index >= 0` selects an entry in the inline-tree blob. //! 3. For each non-negative range, read the 16-byte entry at @@ -78,7 +78,7 @@ pub struct InlineEntry<'a> { /// /// # Bounds /// - /// Real-world Go inline chains are very shallow — typically `< 5` — and + /// Real-world Go inline chains are very shallow - typically `< 5` - and /// the gobin walker caps depth at `32` to keep its cycle-detection /// scratch buffer fixed-size. In practice this field will always fit in /// a `u8`; consumers that store it in a narrower integer (e.g. visus @@ -102,14 +102,14 @@ pub struct InlineTreeIter<'a> { pcdata: Vec<(u32, i32)>, /// Position into `pcdata`. pos: usize, - /// `prev_pc` for the current iteration — start of the next range. + /// `prev_pc` for the current iteration - start of the next range. prev_pc: u32, /// Inline-tree blob bytes (sequence of 16-byte `inlinedCall` records). blob: &'a [u8], } impl<'a> InlineTreeIter<'a> { - /// Construct an iterator that yields nothing — used when prerequisites + /// Construct an iterator that yields nothing - used when prerequisites /// (pclntab, moduledata, funcdata blob) are missing. pub fn empty() -> Self { Self { @@ -175,7 +175,7 @@ impl<'a> InlineTreeIter<'a> { visited_len = visited_len.saturating_add(1); } } else { - // Chain longer than 32 — declare done. Beyond this point we + // Chain longer than 32 - declare done. Beyond this point we // would also exceed any reasonable inlining depth. return depth; } diff --git a/src/structures/itab.rs b/src/structures/itab.rs index 2434e2a..c62c09c 100644 --- a/src/structures/itab.rs +++ b/src/structures/itab.rs @@ -1,4 +1,4 @@ -//! `itablink` decoder — recovers `(interface, concrete type)` pairs. +//! `itablink` decoder - recovers `(interface, concrete type)` pairs. //! //! When the Go linker proves that a concrete type implements an interface, it //! emits an `itab` record carrying both type-descriptor pointers plus a hash @@ -22,7 +22,7 @@ //! ## Why It Matters //! //! Itab pairs let an analyst answer questions like "what implements -//! `io.Reader` in this binary?" — extremely useful when chasing exfiltration +//! `io.Reader` in this binary?" - extremely useful when chasing exfiltration //! paths in malware analysis. use crate::{ @@ -51,7 +51,7 @@ pub struct ItabPair { /// Each [`Iterator::next`] reads one pointer from the underlying itab-array /// (either the `.itablink` section or `moduledata.itablinks`), dereferences /// it through VA→file translation, and parses the [`ItabPair`]. Skips entries -/// that fail to dereference / parse — adversarial input cannot panic the walk. +/// that fail to dereference / parse - adversarial input cannot panic the walk. pub struct ItabIter<'a> { ctx: &'a BinaryContext<'a>, ps: usize, @@ -129,7 +129,7 @@ impl Iterator for ItabIter<'_> { // Advance by the record's true size. `itab.Size()` is // sizeof(itab) (== 4*ptrSize) when `fun[0] == 0`, else // 4*ptrSize + (nmethods-1)*ptrSize. Stop the walk if we - // cannot compute a strictly-positive stride — better to + // cannot compute a strictly-positive stride - better to // truncate than to misalign and emit garbage. let stride = itab_stride(self.ctx, cur, ps, ps_u8)?; let next = cur.checked_add(stride as u64)?; @@ -227,7 +227,7 @@ fn itab_stride(ctx: &BinaryContext<'_>, itab_va: u64, ps: usize, ps_u8: u8) -> O base.checked_add(extra) } -/// Resolve the `Fun[]` method-pointer array of an itab — the concrete-type +/// Resolve the `Fun[]` method-pointer array of an itab - the concrete-type /// implementations bound to each interface method, in interface-method order. /// /// `Fun` sits at `itab_va + 3*ptrSize` and has one `uintptr` per interface diff --git a/src/structures/locate.rs b/src/structures/locate.rs index e406e62..2c6ebe0 100644 --- a/src/structures/locate.rs +++ b/src/structures/locate.rs @@ -7,13 +7,13 @@ //! |-------------------------------|----------------------------------------------| //! | ELF / Mach-O, Go 1.26+ | the dedicated `.go.module` / `__go_module` section | //! | ELF / Mach-O, Go ≤ 1.25 | scan for the `pcHeader` pointer | -//! | PE (every version) | scan — the Go PE linker emits no named Go sections | +//! | PE (every version) | scan - the Go PE linker emits no named Go sections | //! | wasm | scan the reconstructed linear-memory image | //! //! The scan works because `moduledata`'s first field is `pcHeader *pcHeader`, //! which points at the pclntab: a pointer-aligned word equal to the pclntab's //! address, whose surroundings then parse as a plausible moduledata, is the -//! moduledata. Candidates are validated rather than trusted — a pointer to the +//! moduledata. Candidates are validated rather than trusted - a pointer to the //! pclntab can legitimately appear elsewhere in the image. //! //! [`ModuledataLocator`] owns all of that so the type reader and the top-level @@ -71,7 +71,7 @@ impl<'a> ModuledataLocator<'a> { } /// Parse the moduledata out of the dedicated `.go.module` / `__go_module` - /// section (Go 1.26+ ELF and Mach-O). No search needed — the section *is* + /// section (Go 1.26+ ELF and Mach-O). No search needed - the section *is* /// the structure. fn in_section(&self) -> Option { let range = self.ctx.sections().go_module.as_ref()?; @@ -96,7 +96,7 @@ impl<'a> ModuledataLocator<'a> { // The section ranges are file offsets. They index the same bytes the // scan walks for ELF, Mach-O and PE, but for wasm the searched view is // the reconstructed linear-memory image, where a file offset means - // nothing — so wasm scans the image whole. It is the smaller of the + // nothing - so wasm scans the image whole. It is the smaller of the // two anyway, being only the module's initialized data. if self.ctx.format() != BinaryFormat::Wasm { let sections = self.ctx.sections(); @@ -176,7 +176,7 @@ impl<'a> ModuledataLocator<'a> { /// only accepted when the fields around it also hold: the PC range must be /// non-empty, and the structure must be anchored by a field that maps back /// into the image. The legacy (Go 1.5-1.15) layout has no `funcnametab` - /// and — before Go 1.7 — no `types` base, so it is anchored through its + /// and - before Go 1.7 - no `types` base, so it is anchored through its /// always-present `text` boundary instead. fn accept(&self, md: &Moduledata) -> bool { if md.minpc >= md.maxpc { diff --git a/src/structures/maptype.rs b/src/structures/maptype.rs index 85774a5..9e45585 100644 --- a/src/structures/maptype.rs +++ b/src/structures/maptype.rs @@ -7,8 +7,8 @@ //! layouts in all), and //! the meaning of its flag bits changed once more on top of that. Reading one //! layout with another's field list silently shifts every field past the -//! pointer block and — because the extra's size feeds -//! [`crate::structures::descriptor::descriptor_size`] — mislocates the trailing +//! pointer block and - because the extra's size feeds +//! [`crate::structures::descriptor::descriptor_size`] - mislocates the trailing //! `UncommonType`, which then yields a garbage method count. //! //! ## Layouts @@ -35,19 +35,19 @@ //! //! ## Flag bits //! -//! The `flags` word is **not** comparable across the hmap/Swiss boundary — Go +//! The `flags` word is **not** comparable across the hmap/Swiss boundary - Go //! renumbered the bits when it introduced Swiss maps: //! //! | Property | hmap (1.12-1.23) | Swiss (1.24+) | //! |-----------------|------------------|---------------| //! | indirect key | `1 << 0` | `1 << 2` | //! | indirect elem | `1 << 1` | `1 << 3` | -//! | reflexive key | `1 << 2` | — (dropped) | +//! | reflexive key | `1 << 2` | - (dropped) | //! | need key update | `1 << 3` | `1 << 0` | //! | hash might panic| `1 << 4` | `1 << 1` | //! //! Read them through [`MapTypeExtra::flags`], a [`MapFlags`] that normalizes -//! all three encodings — including the pre-1.12 booleans — into `Option` +//! all three encodings - including the pre-1.12 booleans - into `Option` //! per property. [`MapTypeExtra::raw_flags`] keeps the undecoded word for //! callers that want it. //! @@ -63,7 +63,7 @@ use crate::structures::util::{read_u16, read_u32, read_uintptr}; /// Which `abi` map-descriptor layout a binary uses. /// /// Determined from the Go version where one is available, and otherwise from -/// the structural evidence the rest of the binary carries — see +/// the structural evidence the rest of the binary carries - see /// [`MapLayout::infer`]. #[derive(Debug, Clone, Copy, PartialEq, Eq)] pub enum MapLayout { @@ -79,7 +79,7 @@ pub enum MapLayout { Swiss, /// Go 1.27+: Swiss tables carrying explicit key/elem offsets and strides. SwissSplitGroup, - /// Go 1.20-1.25 with no recoverable version string — the moduledata layout + /// Go 1.20-1.25 with no recoverable version string - the moduledata layout /// narrows the window but cannot separate `HmapHasher` (1.20-1.23) from /// `Swiss` (1.24-1.25). Each descriptor is then classified from its own /// bytes by [`MapLayout::resolve_for`]. @@ -102,8 +102,8 @@ impl MapLayout { /// Pick the layout from whatever evidence the binary offers. /// /// The Go version string settles it outright. Without one, the moduledata - /// layout still dates the binary exactly at two boundaries — `V5` is Go - /// 1.27+ and `V4` is Go 1.26 — and the pclntab magic gives a floor: Swiss + /// layout still dates the binary exactly at two boundaries - `V5` is Go + /// 1.27+ and `V4` is Go 1.26 - and the pclntab magic gives a floor: Swiss /// maps arrived in Go 1.24 and therefore imply the Go 1.20 magic, while /// anything below that magic is at most Go 1.19 and so `HmapHasher` /// (the layout in force from 1.14; older magics narrow it no further, and @@ -136,8 +136,8 @@ impl MapLayout { /// Resolve [`Self::Probe`] against one descriptor's extra bytes; every /// other variant returns itself. /// - /// The two candidates in the ambiguous window — `HmapHasher` (Go - /// 1.20-1.23) and `Swiss` (Go 1.24-1.25) — both start with four + /// The two candidates in the ambiguous window - `HmapHasher` (Go + /// 1.20-1.23) and `Swiss` (Go 1.24-1.25) - both start with four /// pointer-sized fields, so they diverge at `extra + 4*ptrSize`, and the /// two readings of that word are disjoint in range: /// @@ -146,8 +146,8 @@ impl MapLayout { /// 128 bytes indirectly (as pointers). So `GroupSize <= 2056` and every /// bit above 15 is zero. /// - **HmapHasher** packs `KeySize u8`, `ValueSize u8`, `BucketSize u16`, - /// `Flags u32` into the same word, and `BucketSize` — which lands in bits - /// 16..31 — is `8 + 8*(KeySize+ValueSize) + ptrSize`, never zero. + /// `Flags u32` into the same word, and `BucketSize` - which lands in bits + /// 16..31 - is `8 + 8*(KeySize+ValueSize) + ptrSize`, never zero. /// /// A non-zero value above bit 15 therefore means `HmapHasher`; anything /// else is `Swiss`. On 32-bit the same reasoning applies to the `u32` at @@ -232,14 +232,14 @@ pub struct MapTypeExtra { /// descriptor. pub group: u64, /// Virtual address of the `hmap` type descriptor - /// ([`MapLayout::HmapWithHmapType`] only — Go 1.11 removed the field). + /// ([`MapLayout::HmapWithHmapType`] only - Go 1.11 removed the field). pub hmap: Option, /// Virtual address of the key-hashing function. `None` for Go 1.11-1.13, /// which had no `hasher` field. pub hasher: Option, /// Size of a slot group in bytes (`GroupSize`). Swiss layouts only. pub group_size: Option, - /// Size of one key/elem slot (`SlotSize`). [`MapLayout::Swiss`] only — + /// Size of one key/elem slot (`SlotSize`). [`MapLayout::Swiss`] only - /// Go 1.27 replaced it with the explicit stride fields below. pub slot_size: Option, /// Offset of the keys array within a group (`KeysOff`). @@ -266,7 +266,7 @@ pub struct MapTypeExtra { pub bucket_size: Option, /// Raw `flags` word. `None` for the pre-1.12 layouts, which encoded the /// same properties as separate booleans. **Bit meanings differ between the - /// hmap and Swiss eras** — prefer [`Self::flags`], which normalizes them. + /// hmap and Swiss eras** - prefer [`Self::flags`], which normalizes them. pub raw_flags: Option, /// Semantic flags, normalized across all three encodings. pub flags: MapFlags, @@ -275,7 +275,7 @@ pub struct MapTypeExtra { impl MapTypeExtra { /// Binary size of the extra for the given pointer size and layout. /// - /// Every layout is `pointer_fields * ptrSize` followed by an 8-byte tail — + /// Every layout is `pointer_fields * ptrSize` followed by an 8-byte tail - /// either four small integers plus four booleans, or `u8 + u8 + u16 + u32`, /// or (Swiss) a `u32` padded out to pointer alignment. That comes to /// `pointers * ps + 8` on 64-bit and `pointers * ps + 8` on 32-bit for the diff --git a/src/structures/mod.rs b/src/structures/mod.rs index a073f1e..199b81e 100644 --- a/src/structures/mod.rs +++ b/src/structures/mod.rs @@ -9,7 +9,7 @@ //! - [`pclntab`] -- The PC/line table: function names, source files, line numbers //! //! [`moduledata`] parses the linker-generated master record that ties the rest -//! together, and [`locate`] finds it — by section where the toolchain emits +//! together, and [`locate`] finds it - by section where the toolchain emits //! one, and by scanning for its `pcHeader` pointer everywhere else. //! //! ## Why These Structures Exist @@ -88,11 +88,11 @@ pub mod wasm; /// Consumers persist this enum into long-lived schemas (database columns, /// structured logs). The contract: /// -/// - **Variants** — append-only. New architectures appear as new variants; +/// - **Variants** - append-only. New architectures appear as new variants; /// existing variants are never renamed or removed. -/// - **`Display` strings** (and [`Self::as_str`]) — fixed forever once +/// - **`Display` strings** (and [`Self::as_str`]) - fixed forever once /// shipped, matching the canonical `GOARCH` values where applicable. -/// - **`Debug` strings** — *not* a stability surface. +/// - **`Debug` strings** - *not* a stability surface. #[derive(Debug, Clone, Copy, PartialEq, Eq)] pub enum Arch { /// x86 32-bit (`minLC=1, ptrSize=4`) @@ -145,11 +145,11 @@ pub enum Arch { /// Consumers persist this enum into long-lived schemas (database columns, /// structured logs). The contract: /// -/// - **Variants** — append-only. New format versions appear as new variants; +/// - **Variants** - append-only. New format versions appear as new variants; /// existing variants are never renamed or removed. -/// - **`Display` strings** (and [`Self::as_str`]) — fixed forever once +/// - **`Display` strings** (and [`Self::as_str`]) - fixed forever once /// shipped. -/// - **`Debug` strings** — *not* a stability surface. +/// - **`Debug` strings** - *not* a stability surface. #[derive(Debug, Clone, Copy, PartialEq, Eq)] pub enum PclntabVersion { /// Go 1.2 through 1.15 (magic `0xFFFFFFFB`). diff --git a/src/structures/moduledata.rs b/src/structures/moduledata.rs index c14e0ae..5112783 100644 --- a/src/structures/moduledata.rs +++ b/src/structures/moduledata.rs @@ -21,8 +21,8 @@ //! //! V5 moved `types`' neighbours, so a wrong guess shifts every field from //! `etypes` onward while still passing a head-only validity check. The -//! out-of-band signals in [`LayoutHints`] are not sufficient on their own — PE -//! carries no `.typelink` section at *any* Go version — so [`Moduledata::parse`] +//! out-of-band signals in [`LayoutHints`] are not sufficient on their own - PE +//! carries no `.typelink` section at *any* Go version - so [`Moduledata::parse`] //! parses both candidates and keeps the one that is internally consistent. use crate::structures::{ @@ -82,7 +82,7 @@ impl TextSect { } } -/// One entry of `moduledata.ptab` (`runtime.ptabEntry`) — an exported symbol of +/// One entry of `moduledata.ptab` (`runtime.ptabEntry`) - an exported symbol of /// a Go plugin. Both fields are offsets relative to `moduledata.types`. #[derive(Debug, Clone, Copy, PartialEq, Eq)] pub struct PtabEntry<'a> { @@ -93,8 +93,8 @@ pub struct PtabEntry<'a> { pub type_offset: i32, } -/// One entry of `moduledata.pkghashes` / `modulehashes` (`runtime.modulehash`) -/// — a per-package ABI hash used to verify plugin / shared-object +/// One entry of `moduledata.pkghashes` / `modulehashes` (`runtime.modulehash`) - +/// a per-package ABI hash used to verify plugin / shared-object /// compatibility at load time. Populated only for plugin / shared builds. #[derive(Debug, Clone, Copy, PartialEq, Eq)] pub struct ModuleHash<'a> { @@ -105,7 +105,7 @@ pub struct ModuleHash<'a> { pub linktime_hash: Option<&'a str>, } -/// A Go `bitvector` (`runtime.bitvector`): `{ n int32; bytedata *byte }` — a +/// A Go `bitvector` (`runtime.bitvector`): `{ n int32; bytedata *byte }` - a /// bit count plus a pointer to the packed bit data. In moduledata these /// describe the GC pointer maps for the data and bss segments. #[derive(Debug, Clone, Copy, Default, PartialEq, Eq)] @@ -161,13 +161,13 @@ pub struct Moduledata { pub text: u64, /// End of text section. pub etext: u64, - /// `[noptrdata, enoptrdata)` — non-pointer initialized data (`.noptrdata`). + /// `[noptrdata, enoptrdata)` - non-pointer initialized data (`.noptrdata`). pub noptrdata: VaRange, - /// `[data, edata)` — pointer-containing initialized data (`.data`). + /// `[data, edata)` - pointer-containing initialized data (`.data`). pub data: VaRange, - /// `[bss, ebss)` — zero-initialized pointer-containing data (`.bss`). + /// `[bss, ebss)` - zero-initialized pointer-containing data (`.bss`). pub bss: VaRange, - /// `[noptrbss, enoptrbss)` — zero-initialized non-pointer data (`.noptrbss`). + /// `[noptrbss, enoptrbss)` - zero-initialized non-pointer data (`.noptrbss`). pub noptrbss: VaRange, /// End VA of the whole module image. pub end: u64, @@ -187,7 +187,7 @@ pub struct Moduledata { pub etypes: u64, /// VA of the start of `.rodata` (Go 1.18+ / V3 / V4 / V5). `None` for V2. pub rodata: Option, - /// VA used as the base for resolving `funcdata[]` offsets — every value + /// VA used as the base for resolving `funcdata[]` offsets - every value /// returned by [`crate::structures::pclntab::ParsedPclntab::funcdata_at`] /// is added to this base to get the funcdata blob's VA. Go 1.18+ / V3 / /// V4 / V5; `None` for V2 binaries (where funcdata used a different @@ -203,7 +203,7 @@ pub struct Moduledata { pub itaboffset: Option, /// Byte length of the itab array (Go 1.27+ / V5 only). `None` for V2-V4. pub itabsize: Option, - /// VA of `epclntab` — one past the end of the pclntab (Go 1.26+ / V4+). + /// VA of `epclntab` - one past the end of the pclntab (Go 1.26+ / V4+). /// `None` pre-1.26. pub epclntab: Option, /// inittasks slice: `[]*initTask`, the linker-built list of package @@ -212,32 +212,32 @@ pub struct Moduledata { /// `textsectmap` slice (one `textsect` per text section). Length > 1 only /// for large binaries the linker split across multiple text sections. pub textsectmap: GoSlice, - /// `ptab` slice (`[]ptabEntry`) — exported plugin symbols. Non-empty only + /// `ptab` slice (`[]ptabEntry`) - exported plugin symbols. Non-empty only /// for `-buildmode=plugin`. pub ptab: GoSlice, /// `pluginpath` string header. Non-empty only for `-buildmode=plugin`. pub pluginpath: GoStr, - /// `pkghashes` slice (`[]modulehash`) — per-package ABI hashes used to + /// `pkghashes` slice (`[]modulehash`) - per-package ABI hashes used to /// verify plugin/shared compatibility. Non-empty only for plugin/shared. pub pkghashes: GoSlice, /// `modulename` string header. Set for plugins / shared libraries; empty /// for an ordinary executable. pub modulename: GoStr, - /// `modulehashes` slice (`[]modulehash`) — dependency ABI hashes for + /// `modulehashes` slice (`[]modulehash`) - dependency ABI hashes for /// plugin/shared compatibility checks. pub modulehashes: GoSlice, - /// `hasmain` flag — this module contains the program's `main` (true for + /// `hasmain` flag - this module contains the program's `main` (true for /// the main executable, false for plugins / shared libraries). Best-effort /// tail read; `false` if the tail was truncated. pub has_main: bool, - /// `bad` flag — the runtime marks a module that failed to load and should + /// `bad` flag - the runtime marks a module that failed to load and should /// be ignored. Best-effort tail read. pub bad: bool, - /// `gcdatamask` bitvector — GC pointer map for the data segment. + /// `gcdatamask` bitvector - GC pointer map for the data segment. pub gcdatamask: Bitvector, - /// `gcbssmask` bitvector — GC pointer map for the bss segment. + /// `gcbssmask` bitvector - GC pointer map for the bss segment. pub gcbssmask: Bitvector, - /// VA of the runtime `typemap` (`map[typeOff]*_type`) — cross-module type + /// VA of the runtime `typemap` (`map[typeOff]*_type`) - cross-module type /// deduplication map, populated at load time. `0` if absent. pub typemap: u64, /// VA of the `next` moduledata in the linked list, or `0` for the last @@ -287,8 +287,8 @@ pub struct LayoutHints { /// binaries whose version was stripped or obfuscated away. pub go_minor: Option, /// Whether a `.typelink` / `__typelink` section was found. Its *absence* - /// is weak evidence of Go 1.27+ — PE, wasm, and RELRO ELF links spelled - /// `.data.rel.ro.typelink` have no such section at any version — so it is + /// is weak evidence of Go 1.27+ - PE, wasm, and RELRO ELF links spelled + /// `.data.rel.ro.typelink` have no such section at any version - so it is /// never used alone. See [`Moduledata::parse`]. pub has_typelink_section: bool, /// Whether a `.go.type` / `__go_type` section was found. This one *is* a @@ -339,15 +339,15 @@ impl Moduledata { /// V5 moved `types`' neighbours around (`+typedesclen`, `+itaboffset`, /// `+itabsize`, `-typelinks`, `-itablinks`), so guessing wrong shifts every /// field from `etypes` onward and yields a moduledata that still passes a - /// head-only validity check — `minpc`/`maxpc`/`funcnametab` are ahead of - /// the divergence — while silently reporting an empty types region, no + /// head-only validity check - `minpc`/`maxpc`/`funcnametab` are ahead of + /// the divergence - while silently reporting an empty types region, no /// itabs, no init tasks and `has_main == false`. /// /// The hints alone cannot settle it: PE never emits a `.typelink` section /// at any Go version, so on a PE binary whose version string was scrubbed /// the only remaining signal points the wrong way. Instead of trusting the /// hints, this parses *both* candidate layouts (hint-preferred first) and - /// returns the first one that is internally self-consistent — see + /// returns the first one that is internally self-consistent - see /// [`Moduledata::layout_self_consistent`]. Only if neither validates does /// the hint-preferred layout win, so callers still get the head fields. pub fn parse(data: &[u8], ps: u8, hints: LayoutHints) -> Option { @@ -372,7 +372,7 @@ impl Moduledata { Self::parse_modern(data, ps, hints, prefer_v5) } - /// Whether the parsed layout is internally consistent — the arbiter used by + /// Whether the parsed layout is internally consistent - the arbiter used by /// [`Self::parse`] to decide whether it guessed V5 correctly. /// /// Checks only relationships that hold in *every* real Go image, so a @@ -562,7 +562,7 @@ impl Moduledata { let textsectmap = GoSlice::parse(data, off, ps).unwrap_or_default(); off = advance(off, slice_sz)?; - // typelinks, itablinks (slices) — removed in V5. + // typelinks, itablinks (slices) - removed in V5. let (typelinks, itablinks) = if v5 { (None, None) } else { @@ -581,7 +581,7 @@ impl Moduledata { let pkghashes = GoSlice::parse(data, off, ps).unwrap_or_default(); off = advance(off, slice_sz)?; - // inittasks []*initTask (Go 1.21+) — best-effort tail read; a malformed + // inittasks []*initTask (Go 1.21+) - best-effort tail read; a malformed // tail collapses to `None` rather than failing the whole parse. let inittasks = if has_inittasks { let it = GoSlice::parse(data, off, ps); @@ -591,7 +591,7 @@ impl Moduledata { None }; - // Tail (all best-effort — a truncated moduledata must not fail the + // Tail (all best-effort - a truncated moduledata must not fail the // whole parse, since the parser needs nothing past this point): // modulename (string), modulehashes (slice), hasmain (uint8), then a // version-divergent block. In Go 1.24+ `bad` moved to right after @@ -714,7 +714,7 @@ impl Moduledata { /// Source: `src/runtime/symtab.go` `moduledata` (Go 1.5-1.15). fn parse_go12_legacy(data: &[u8], ps: u8, go_version_minor: Option) -> Option { // `moduledata` was introduced in Go 1.5. Go 1.2-1.4 have none, so when - // the version is known to predate 1.5 we refuse to parse — otherwise a + // the version is known to predate 1.5 we refuse to parse - otherwise a // pclntab-address pointer elsewhere in the image could be mistaken for // a (false) moduledata. When the version is unknown we still attempt it // and rely on the locator's structural validation. @@ -830,7 +830,7 @@ impl Moduledata { (GoSlice::default(), GoStr::default(), GoSlice::default()) }; - // Tail (best-effort — a truncated moduledata must not fail the parse): + // Tail (best-effort - a truncated moduledata must not fail the parse): // modulename, modulehashes, [hasmain], gcdatamask, gcbssmask, // [typemap], [bad], next. let modulename = GoStr::parse(data, off, ps).unwrap_or_default(); @@ -952,8 +952,8 @@ fn looks_like_slice_header(data: &[u8], off: usize, ps: u8) -> bool { mod tests { use super::*; - /// Hints for a binary with neither a `.typelink` nor a `.go.type` section - /// — the shape a PE, a wasm module, or a stripped RELRO ELF presents. + /// Hints for a binary with neither a `.typelink` nor a `.go.type` section - + /// the shape a PE, a wasm module, or a stripped RELRO ELF presents. fn hints(pclntab_version: PclntabVersion, go_minor: Option) -> LayoutHints { LayoutHints { pclntab_version, @@ -1132,7 +1132,7 @@ mod tests { // Regression guard: PE binaries never carry a `.typelink` section, so // for a version-scrubbed PE the only hint points at V5. Reading a // pre-V5 moduledata with the V5 layout collapses the types span while - // leaving a huge `typedesclen`, which the consistency check rejects — + // leaving a huge `typedesclen`, which the consistency check rejects - // so the parser must fall back to V3 rather than return the garbage. let data = synthetic_v3(); let md = Moduledata::parse(&data, 8, hints(PclntabVersion::Go120, None)) @@ -1180,7 +1180,7 @@ mod tests { #[test] fn v5_is_never_chosen_below_the_go120_magic() { - // V5 requires covctrs, which the Go118 magic rules out — so even with + // V5 requires covctrs, which the Go118 magic rules out - so even with // every V5 hint set the layout must stay pre-V5. let data = synthetic_v5(); let md = Moduledata::parse( diff --git a/src/structures/pclntab.rs b/src/structures/pclntab.rs index 9d42191..35151c7 100644 --- a/src/structures/pclntab.rs +++ b/src/structures/pclntab.rs @@ -136,13 +136,13 @@ const MAGICS: &[([u8; 4], PclntabVersion)] = &[ /// = the `pcHeader` magic bytes). The lifetime `'a` borrows from the address /// space the parser ran on: /// -/// - For ELF / Mach-O / PE this is the input file bytes — zero-copy access. +/// - For ELF / Mach-O / PE this is the input file bytes - zero-copy access. /// - For Wasm this is the reconstructed linear-memory image owned by /// [`crate::formats::BinaryContext`]; in that case the borrow lives only as /// long as the `&BinaryContext` it was derived from. Cache the scalar /// metadata via [`Self::meta`] if you need to re-attach later. /// -/// Cheap to copy — every field is `Copy` (`&'a [u8]` and integers). +/// Cheap to copy - every field is `Copy` (`&'a [u8]` and integers). #[derive(Debug, Clone, Copy)] pub struct ParsedPclntab<'a> { /// The entire pclntab section data, starting at the pcHeader. @@ -186,7 +186,7 @@ pub struct ParsedPclntab<'a> { /// longer stored… Code should use the moduledata text field instead."*). /// /// `None` for Go 1.16-1.17, which lack the field, and for Go 1.26+, where - /// the slot reads as zero — a zero is reported as absent rather than as + /// the slot reads as zero - a zero is reported as absent rather than as /// address `0`, so callers do not silently rebase every function entry /// against the bottom of the address space. Mirrors `moduledata.text`. pub header_text_start: Option, @@ -235,13 +235,13 @@ impl PclntabMeta { /// /// `address_data` is the full address-space view the metadata was /// originally derived from (for ELF/Mach-O/PE, the input file bytes; for - /// wasm, the reconstructed linear-memory image — typically obtained via + /// wasm, the reconstructed linear-memory image - typically obtained via /// [`crate::formats::BinaryContext::structure_search_data`]). This /// helper slices it at [`Self::offset`] so the returned struct's /// `data[0]` is the pcHeader's first byte, matching the layout /// `pclntab::parse` produces. /// - /// Returns `None` if `address_data` does not cover `self.offset` — + /// Returns `None` if `address_data` does not cover `self.offset` - /// would only happen if the buffer the metadata came from is not the one /// passed in. pub fn attach<'a>(&self, address_data: &'a [u8]) -> Option> { @@ -291,7 +291,7 @@ impl<'a> ParsedPclntab<'a> { /// See [`Arch`] for the mapping table and caveats. Several /// `(minLC, ptrSize)` combinations are ambiguous; in particular /// `(1, 8)` matches both `Arch::X86_64` and `Arch::Wasm`. This accessor - /// always reports `Arch::X86_64` for that combination — use + /// always reports `Arch::X86_64` for that combination - use /// [`crate::GoBinary::arch`] for the format-disambiguated result. pub fn arch(&self) -> Arch { match (self.min_lc, self.ptr_size) { @@ -346,7 +346,7 @@ impl<'a> ParsedPclntab<'a> { /// measured from. /// /// In Go 1.16+, `funcoff` is relative to the functab section start - /// ([`Self::functab_offset`]) — the `_func` structs sit there, after the + /// ([`Self::functab_offset`]) - the `_func` structs sit there, after the /// `(nfunc+1)` entry pairs. In Go 1.2-1.15 there is no separate functab /// section: `funcoff` is an offset from the pcHeader itself (`data[0]`), so /// the base is `0`. @@ -438,7 +438,7 @@ impl<'a> ParsedPclntab<'a> { /// Read the i-th `funcdata[]` offset for the given function. The value /// is a `u32` offset into `moduledata.gofunc`; `0xFFFFFFFF` (`^uint32(0)`) - /// is the sentinel for "no funcdata at this index" — callers should treat + /// is the sentinel for "no funcdata at this index" - callers should treat /// it as `None`. /// /// Returns `None` if `i >= func.nfuncdata` or the read goes out of bounds. @@ -522,7 +522,7 @@ impl<'a> ParsedPclntab<'a> { /// /// pcfile entries whose index doesn't resolve through the cutab are /// skipped (consistent with [`Self::decode_pcfile_paths`]). pcln events - /// before the first pcfile transition are also skipped — they would + /// before the first pcfile transition are also skipped - they would /// have no file to attribute to. pub fn decode_pcln_with_files<'pcl>(&'pcl self, func: &FuncData) -> PcLineFileIter<'pcl, 'a> { PcLineFileIter { @@ -927,7 +927,7 @@ impl<'a> Iterator for PcFilePathIter<'_, 'a> { /// Yields `(pc_offset, line, file_path)`. The file table only emits a /// transition when the active source file *changes*, so the joined iterator /// carries the latest transition forward across `pcln` events. Entries whose -/// file index doesn't resolve through the cutab are skipped silently — same +/// file index doesn't resolve through the cutab are skipped silently - same /// rule as [`PcFilePathIter`]. /// /// This replaces the hand-rolled "walk pcfile and pcln in lockstep" state @@ -952,7 +952,7 @@ impl<'a> Iterator for PcLineFileIter<'_, 'a> { // Advance the pcfile cursor to the latest transition with pc ≤ current. // pcfile transitions are emitted in monotonically increasing PC order, - // and we walk pcln in the same order — so the cursor only ever moves + // and we walk pcln in the same order - so the cursor only ever moves // forward. let mut next_idx = match self.cursor { Some(c) => c.saturating_add(1), @@ -971,7 +971,7 @@ impl<'a> Iterator for PcLineFileIter<'_, 'a> { let cursor = match self.cursor { Some(c) => c, - // No pcfile transition yet ≤ current pc — try the next pcln event. + // No pcfile transition yet ≤ current pc - try the next pcln event. None => continue, }; let (_, file_idx) = match self.pcfile.get(cursor) { @@ -1067,19 +1067,19 @@ impl FuncData { /// /// Uses a layered detection strategy, from cheapest/most reliable to most expensive: /// -/// 1. **Known section + magic** — If the binary has a `.gopclntab` section, validate +/// 1. **Known section + magic** - If the binary has a `.gopclntab` section, validate /// its start against the known magic bytes. This is the fastest and most reliable path. -/// 2. **Full magic scan** — Scan the entire binary at 4-byte aligned offsets for one +/// 2. **Full magic scan** - Scan the entire binary at 4-byte aligned offsets for one /// of the four known magic values, then validate the full header. -/// 3. **Relaxed header scan** — If magic bytes were wiped (common in malware), scan +/// 3. **Relaxed header scan** - If magic bytes were wiped (common in malware), scan /// for the pcHeader structural pattern without requiring magic. Validates /// `pad1==0, pad2==0, minLC∈{1,2,4}, ptrSize∈{4,8}` plus sub-table offset /// monotonicity and funcname spot-checking. (Strategy A from RESEARCH.md §1.5) -/// 4. **moduledata pointer chain** — Scan data sections for pointer-aligned values +/// 4. **moduledata pointer chain** - Scan data sections for pointer-aligned values /// that point into the `.gopclntab` section range, then validate the target with /// relaxed header validation. This mirrors how the Go runtime itself finds the /// pclntab via `runtime.firstmoduledata.pcHeader`. (Strategy B from RESEARCH.md §1.5) -/// 5. **functab monotonicity** — Scan read-only sections for long arrays of +/// 5. **functab monotonicity** - Scan read-only sections for long arrays of /// `(u32, u32)` pairs with strictly monotonically increasing first elements. /// Work backwards to find the pcHeader. (Strategy C from RESEARCH.md §1.5) pub fn parse<'a>(ctx: &'a BinaryContext<'a>) -> Option> { @@ -1169,7 +1169,7 @@ fn scan_for_magic_strided(data: &[u8], stride: usize) -> Option Option> { if data.len() < 8 { return None; @@ -1298,7 +1298,7 @@ fn scan_via_moduledata<'a>(ctx: &BinaryContext<'a>) -> Option> }; // Must point to the start of the gopclntab section (pcHeader is at the beginning) - // Allow a small tolerance — the pointer should be within the first 64 bytes + // Allow a small tolerance - the pointer should be within the first 64 bytes if candidate_va >= pclntab_va && candidate_va < header_window_end && candidate_va < pclntab_va_end @@ -1520,14 +1520,14 @@ fn recover_header_from_functab<'a>( /// Parse a pcHeader at the start of `data` with a known version. /// -/// Validates field ranges (nfunc, nfiles, offsets) but does NOT check magic bytes -/// — the caller is responsible for version determination. +/// Validates field ranges (nfunc, nfiles, offsets) but does NOT check magic bytes - +/// the caller is responsible for version determination. fn parse_header( data: &[u8], base_offset: usize, version: PclntabVersion, ) -> Option> { - // The Go 1.2-1.15 pclntab has no structured pcHeader — a different parser. + // The Go 1.2-1.15 pclntab has no structured pcHeader - a different parser. if version == PclntabVersion::Go12 { return parse_header_go12(data, base_offset); } @@ -1592,8 +1592,8 @@ fn parse_header( // The pcHeader gained a `textStart` field at index 2 in Go 1.18; Go // 1.16-1.17 (off_base == 2) do not have it, and Go 1.26+ zeroed it out // (see `ParsedPclntab::header_text_start`). `runtime.text` is never 0 in a - // real image — even wasm, whose PCs start at 0, carries the value in the - // moduledata rather than here — so a zero means "not recorded". + // real image - even wasm, whose PCs start at 0, carries the value in the + // moduledata rather than here - so a zero means "not recorded". let header_text_start = if off_base > 2 { let off = advance_n(8, 2, ps)?; read_uintptr(data, off, ptr_size).filter(|&va| va != 0) @@ -1624,7 +1624,7 @@ fn parse_header( /// Unlike Go 1.16+, the legacy format has no structured `pcHeader`: the 8-byte /// prefix (`magic`, two pad bytes, `minLC`, `ptrSize`) is followed immediately /// by a pointer-sized `nfunctab` and the functab itself. There is no separate -/// `funcnametab`, `cutab`, or `pctab` — function names and the `pcsp`/`pcfile`/ +/// `funcnametab`, `cutab`, or `pctab` - function names and the `pcsp`/`pcfile`/ /// `pcln` tables are addressed by offsets relative to the pcHeader (`data[0]`), /// so [`ParsedPclntab::funcname_offset`] and [`ParsedPclntab::pctab_offset`] /// are both `0`. The filetab is located by a `u32` stored immediately after the @@ -1726,7 +1726,7 @@ fn parse_header_go12(data: &[u8], base_offset: usize) -> Option, /// `startLine` offset (Go 1.20+), else `None` (read as 0). @@ -1908,7 +1908,7 @@ mod tests { #[test] fn test_strategy_a_relaxed_header_zeroed_magic() { - // Build a valid pclntab, then zero the magic — relaxed scan should still find it + // Build a valid pclntab, then zero the magic - relaxed scan should still find it let mut data = build_synthetic_pclntab([0xf1, 0xff, 0xff, 0xff]); data[0..4].copy_from_slice(&[0x00, 0x00, 0x00, 0x00]); // wipe magic diff --git a/src/structures/strings.rs b/src/structures/strings.rs index c5017db..49ecff3 100644 --- a/src/structures/strings.rs +++ b/src/structures/strings.rs @@ -1,6 +1,6 @@ //! Go-style string literal scanner. //! -//! Go strings are stored as `(ptr, len)` headers — *not* NUL-terminated — +//! Go strings are stored as `(ptr, len)` headers - *not* NUL-terminated - //! with the actual UTF-8 bytes living in `.rodata` (or equivalent for //! non-ELF). A generic strings extractor either misses them entirely or //! splits them at internal NULs. This module provides a precise scanner: @@ -11,7 +11,7 @@ //! //! ## Heuristics //! -//! False positives are inherent to a `(u64, u64)` scan — any random pair +//! False positives are inherent to a `(u64, u64)` scan - any random pair //! that happens to look like `(in-segment ptr, plausible len)` and points //! to UTF-8 bytes will match. We minimize them by: //! @@ -26,7 +26,7 @@ //! Consumers decide whether to interpret bytes as text via [`GoString::as_bytes`], //! [`GoString::try_as_str`], or the convenience [`GoString::as_str`]. //! -//! Duplicate yields are *not* filtered — a string referenced from N +//! Duplicate yields are *not* filtered - a string referenced from N //! different positions yields N times. Consumers that want unique results //! can `.collect::>()`. @@ -61,7 +61,7 @@ pub struct GoString<'a> { impl<'a> GoString<'a> { /// Raw bytes of the string, borrowed from the binary. /// - /// Always succeeds — this is the canonical view for callers that want + /// Always succeeds - this is the canonical view for callers that want /// to handle arbitrary byte content (hex-dump, base64-encode, MinHash, /// etc.). Unlike [`Self::as_str`] / [`Self::try_as_str`], no UTF-8 /// validation is performed. @@ -121,7 +121,7 @@ impl<'a> Iterator for GoStringIter<'a> { return None; } let ps_u8 = u8::try_from(ps).ok()?; - // Walk the address space the runtime would see — for wasm the + // Walk the address space the runtime would see - for wasm the // reconstructed linear-memory image, otherwise the file bytes. let data = self.ctx.structure_search_data(); @@ -155,7 +155,7 @@ impl<'a> Iterator for GoStringIter<'a> { Some(b) => b, None => continue, }; - // No UTF-8 filter here — see the module-level docs. Callers that + // No UTF-8 filter here - see the module-level docs. Callers that // need text use `GoString::try_as_str` / `as_str`; callers that // want raw rodata bytes (malware payloads, MinHash signal) use // `GoString::as_bytes`. @@ -181,7 +181,7 @@ pub fn extract_iter<'a>( let (text_start, text_end) = match moduledata { Some(m) => (m.text, m.etext), // Without moduledata we can't filter text pointers; everything is a - // candidate. That's acceptable — the UTF-8 + length filters still + // candidate. That's acceptable - the UTF-8 + length filters still // cut most noise. None => (0, 0), }; diff --git a/src/structures/types.rs b/src/structures/types.rs index 4fa938d..9cefdd2 100644 --- a/src/structures/types.rs +++ b/src/structures/types.rs @@ -11,7 +11,7 @@ //! //! 2. **Descriptor-walking path** (Go 1.27+, which has no typelink table): Walk //! from `moduledata.types + PtrSize` to `moduledata.types + typedesclen`, -//! advancing by each type's `DescriptorSize` with pointer alignment — the +//! advancing by each type's `DescriptorSize` with pointer alignment - the //! same algorithm as the Go runtime's `moduleTypelinks()`. Nothing past //! `typedesclen` is walkable: the linker groups every `type:`-prefixed //! read-only symbol under one carrier and sorts the non-typelink remainder @@ -28,7 +28,7 @@ //! - Type descriptors: `src/internal/abi/type.go` //! - Type walking: `src/runtime/type.go` (`moduleTypelinks`) //! - Type-section layout: `src/cmd/link/internal/ld/data.go` (`dodataSect`, -//! `sym.STYPE` case) — which also records `typedesclen` and `itaboffset` +//! `sym.STYPE` case) - which also records `typedesclen` and `itaboffset` //! - Moduledata: `src/runtime/symtab.go` use std::collections::{HashSet, VecDeque}; @@ -105,7 +105,7 @@ pub struct GoType<'a> { /// regular-memory, gc-mask-on-demand, direct-iface). `is_named` / /// `is_exported` are derived from it; this exposes the unparsed value. pub tflag: u8, - /// `PtrToThis` — `TypeOff` (offset from `moduledata.types`) of the + /// `PtrToThis` - `TypeOff` (offset from `moduledata.types`) of the /// pointer-to-this (`*T`) type descriptor, or `0` if the linker emitted /// none. pub ptr_to_this: i32, @@ -130,7 +130,7 @@ pub struct GoType<'a> { pub exported_method_count: u16, /// Kind-specific type details parsed from the type descriptor's extra fields. pub detail: TypeDetail<'a>, - /// Resolved method list for this type (concrete-type methods only — + /// Resolved method list for this type (concrete-type methods only - /// interface methods live on [`TypeDetail::Interface`]). Empty if /// [`Self::has_uncommon`] is `false`. pub methods: Vec>, @@ -141,7 +141,7 @@ pub struct GoType<'a> { /// `"[]uint8"`, `"map[string]int"`). /// /// `name` is `None` when the descriptor could not be located or carries no -/// name — callers keep the raw `va` and avoid fabricating a label. +/// name - callers keep the raw `va` and avoid fabricating a label. #[derive(Debug, Clone, Copy, PartialEq, Eq)] pub struct TypeRef<'a> { /// Virtual address of the referenced type descriptor. @@ -158,7 +158,7 @@ pub struct MethodEntry<'a> { /// Type-descriptor offset relative to `moduledata.types` (`mtyp` field). pub type_descriptor_offset: i32, /// Resolved name of the method's signature-type descriptor, if it carries - /// one. Usually `None`: Go func types are unnamed, so the *name* is empty — + /// one. Usually `None`: Go func types are unnamed, so the *name* is empty - /// resolve [`Self::type_descriptor_offset`] to a [`TypeDetail::Func`] /// (whose params now carry names) for the full signature. pub type_name: Option<&'a str>, @@ -217,11 +217,11 @@ pub struct InterfaceMethod<'a> { /// Consumers persist this enum's *kind tag* into long-lived schemas /// (database columns, structured logs). The contract: /// -/// - **Variants** — append-only. New kinds appear as new variants; existing +/// - **Variants** - append-only. New kinds appear as new variants; existing /// variants are never renamed or removed. -/// - **[`Self::kind_str`]** — fixed forever once shipped; treat the +/// - **[`Self::kind_str`]** - fixed forever once shipped; treat the /// returned strings as serialization keys. -/// - **`Debug` strings** — *not* a stability surface. +/// - **`Debug` strings** - *not* a stability surface. #[derive(Debug, Clone)] pub enum TypeDetail<'a> { /// No extra detail (scalar types, string, unsafe.Pointer). @@ -572,7 +572,7 @@ impl std::fmt::Display for TypeKind { /// Streaming iterator over [`GoType`]s extracted from a binary. /// /// Backed by [`extract_types_iter`]. Each [`Iterator::next`] call parses one -/// `abi.Type` lazily — no `Vec` is allocated up front. Skips any descriptor +/// `abi.Type` lazily - no `Vec` is allocated up front. Skips any descriptor /// that fails to parse (adversarial input cannot panic the iteration). pub struct TypeIter<'a> { ctx: &'a BinaryContext<'a>, @@ -586,7 +586,7 @@ enum TypeIterStrategy<'a> { /// Iterate `int32` offsets from a typelink array. Typelinks { tl_data: &'a [u8], pos: usize }, /// Walk a contiguous descriptor region `[td, end_va)`, advancing by - /// `DescriptorSize` with pointer alignment — the same stepper + /// `DescriptorSize` with pointer alignment - the same stepper /// `runtime.moduleTypelinks` uses on Go 1.27+. Walk { /// VA of the next descriptor to parse. @@ -603,7 +603,7 @@ enum TypeIterStrategy<'a> { } impl<'a> TypeIter<'a> { - /// An iterator that yields nothing — the result when a binary has no + /// An iterator that yields nothing - the result when a binary has no /// moduledata, no VA mapping, or no types region. pub fn empty(ctx: &'a BinaryContext<'a>) -> Self { Self { @@ -648,7 +648,7 @@ impl<'a> Iterator for TypeIter<'a> { { return Some(go_type); } - // Failed to parse this entry — fall through to the next. + // Failed to parse this entry - fall through to the next. } None } @@ -723,7 +723,7 @@ impl<'a> Iterator for TypeIter<'a> { } /// Construct a streaming iterator over the binary's **reflection-visible** -/// type descriptors — the set the `typelink` table used to name. +/// type descriptors - the set the `typelink` table used to name. /// /// The constructor performs moduledata discovery up front (cheap on ELF / /// Mach-O, scan-based on PE) so each [`Iterator::next`] call does only the @@ -810,8 +810,8 @@ const WALK_SKIP_BUDGET: u32 = 64; /// On V5 the region is bounded by `moduledata.typedesclen`, exactly as /// `runtime.moduleTypelinks` reads it; the descriptors past that bound are /// non-typelink types and then itabs, neither of which belongs in the -/// typelink enumeration. Pre-V5 binaries have no such bound — they only reach -/// this walk when both typelink tables are unavailable — so the whole types +/// typelink enumeration. Pre-V5 binaries have no such bound - they only reach +/// this walk when both typelink tables are unavailable - so the whole types /// region is used. /// /// The `ptrSize` skip at the head is the slot the linker reserves so that no @@ -822,7 +822,7 @@ const WALK_SKIP_BUDGET: u32 = 64; /// the non-typelink remainder by size, so `[typedesclen, itaboffset)` is a mix /// of non-typelink descriptors and `type:.namedata.*` blobs with no recorded /// boundary between them. Those descriptors are reachable only by following -/// references — see [`extract_all_types`]. +/// references - see [`extract_all_types`]. fn typelink_walk_range(md: &Moduledata, ptr_size: u8) -> Option<(u64, u64)> { let start = md.types.checked_add(u64::from(ptr_size))?; let end = match md.typedesclen { @@ -877,7 +877,7 @@ pub fn type_at_va<'a>( /// Transitively enumerate every type reachable from `seeds`. /// /// BFS over type-descriptor virtual addresses: each popped VA is parsed -/// independently (robust — no reliance on descriptor sizing), and all the +/// independently (robust - no reliance on descriptor sizing), and all the /// types it references are enqueued. Reaches types absent from the seed set /// (typically `typelink`), e.g. a struct used only as a pointer's element. pub fn extract_all_types<'a>( @@ -987,7 +987,7 @@ fn resolve_name_at_va<'a>( /// Resolve a referenced type descriptor at `type_va` to its display name. /// /// Parses the `abi.Type` at the target VA and decodes its `Str` (a NameOff -/// relative to `types_base_va`). One level deep only — Go already stores +/// relative to `types_base_va`). One level deep only - Go already stores /// constructed names like `[]uint8` / `map[string]int` in `Str` for composite /// types, so this yields a usable label without reimplementing the runtime's /// recursive type formatter. Returns `None` (rather than guessing) when the @@ -1232,12 +1232,12 @@ fn build_go_type<'a>( /// Layout (after the embedded `abi.Type`): /// - 4 bytes: `FuncTypeExtra` (`InCount` u16, `OutCount` u16) /// - Padding to pointer-size alignment -/// - `UncommonType` (16 bytes) when the type carries one — the parameter array +/// - `UncommonType` (16 bytes) when the type carries one - the parameter array /// follows it (see `abi.FuncType.InSlice`) /// - `(in_count + out_count) * ps` bytes: `*Type` pointers /// /// Returns `(inputs, outputs)`. Lengths may be shorter than the requested -/// counts on truncated input — callers should treat that as malformed. +/// counts on truncated input - callers should treat that as malformed. #[allow(clippy::too_many_arguments)] fn read_func_params<'a>( type_data: &'a [u8], @@ -1256,7 +1256,7 @@ fn read_func_params<'a>( return (Vec::new(), Vec::new()); } // Params start after the (ptr-aligned) `funcType` struct, then after the - // `UncommonType` when present — Go places the inline parameter array *after* + // `UncommonType` when present - Go places the inline parameter array *after* // the uncommon block (`abi.FuncType.InSlice`), so skipping it here is what // keeps the parameter VAs (and the types they reach) correct. let uncommon_sz = if has_uncommon { UncommonType::SIZE } else { 0 }; @@ -1331,7 +1331,7 @@ fn resolve_concrete_methods<'a>( }; // A concrete-type method always has a name. An empty / unresolved name // means `mcount` over-ran the real method array and the loop is now - // reading unrelated bytes — this happens when a stray reference is + // reading unrelated bytes - this happens when a stray reference is // mis-parsed as a type descriptor and its `UncommonType.mcount` is // garbage (e.g. `0xFFFF` read from padding). Stop rather than fabricate // thousands of empty methods. `mcount` cannot be bounded by a VA range diff --git a/src/structures/util.rs b/src/structures/util.rs index 6d8b682..421b8ef 100644 --- a/src/structures/util.rs +++ b/src/structures/util.rs @@ -50,7 +50,7 @@ pub(crate) fn read_u16(data: &[u8], offset: usize) -> Option { /// Advance a cursor `off` by `by` bytes, returning `None` on overflow. /// -/// Idiomatic shorthand for `off.checked_add(by)` — used by sequential parsers +/// Idiomatic shorthand for `off.checked_add(by)` - used by sequential parsers /// like [`crate::structures::moduledata::Moduledata::parse`] that walk /// fixed-layout records field by field. #[inline] diff --git a/src/structures/wasm.rs b/src/structures/wasm.rs index 11429cb..77553b4 100644 --- a/src/structures/wasm.rs +++ b/src/structures/wasm.rs @@ -6,7 +6,7 @@ //! binary: //! //! - [`walk`] enumerates top-level sections, surfacing custom-section names -//! (`go:buildid`, `producers`, `name` — per `src/cmd/link/internal/wasm/asm.go`) +//! (`go:buildid`, `producers`, `name` - per `src/cmd/link/internal/wasm/asm.go`) //! and Data-section payload bounds. The Go linker does *not* emit a //! dedicated `.gopclntab` custom section for wasm; pclntab + buildinfo //! live inside the Data-section linear-memory payload, alongside the rest @@ -23,7 +23,7 @@ //! parsers address them by their runtime VA the same way they do on //! ELF/Mach-O/PE. //! -//! This module is intentionally narrow — it does not understand wasm +//! This module is intentionally narrow - it does not understand wasm //! function types, imports, exports, instructions, or linking metadata. //! Anything more is the caller's job. //! @@ -39,7 +39,7 @@ pub struct WasmSection<'a> { /// Custom-section name when `id == 0`, else `None`. pub name: Option<&'a str>, /// File-offset range covering just the section's payload (excludes the - /// id byte, the LEB128 length prefix, and — for custom sections — the + /// id byte, the LEB128 length prefix, and - for custom sections - the /// name header). pub payload_offset: usize, /// Length of the payload range. @@ -79,7 +79,7 @@ impl<'a> Iterator for WasmSectionIter<'a> { let size = usize::try_from(size).ok()?; let payload_end = after_len.checked_add(size)?; if payload_end > self.data.len() { - // Malformed section — abort the walk. + // Malformed section - abort the walk. self.pos = self.data.len(); return None; } @@ -186,7 +186,7 @@ pub fn data_segments(data: &[u8]) -> Vec { u64::from(val as u32) } _ => { - // Unsupported offset expression — bail rather than + // Unsupported offset expression - bail rather than // silently miscompute. return out; } @@ -230,7 +230,7 @@ pub fn data_segments(data: &[u8]) -> Vec { /// Build a linear-memory image from wasm data segments, copying each /// segment's bytes to its target offset and zero-filling gaps. /// -/// Returns `None` if the resulting image would exceed `max_size_bytes` — +/// Returns `None` if the resulting image would exceed `max_size_bytes` - /// adversarial input could request gigabytes of zero-fill, so callers gate /// this behind a sanity cap. The returned vector's length is exactly the /// largest `mem_offset + size` across all segments. @@ -264,7 +264,7 @@ pub fn build_linear_memory_image(data: &[u8], max_size_bytes: usize) -> Option Option<(u32, usize)> { let mut result: u32 = 0; diff --git a/tests/integration.rs b/tests/integration.rs index cc06eb9..7492250 100644 --- a/tests/integration.rs +++ b/tests/integration.rs @@ -485,7 +485,7 @@ mod strings_and_inline { #[test] fn strings_iter_works_on_stripped_binaries() { - // Strings live in rodata, not in DWARF or symbol tables — should + // Strings live in rodata, not in DWARF or symbol tables - should // survive `-ldflags='-s -w'` unchanged. let normal = { let d = load(BASIC_NORMAL); @@ -497,7 +497,7 @@ mod strings_and_inline { let bin = GoBinary::parse(&d).unwrap(); bin.strings().count() }; - // Should be very close — within 5% (stripping shouldn't materially affect rodata). + // Should be very close - within 5% (stripping shouldn't materially affect rodata). let diff = (normal as i64 - stripped as i64).abs(); let bound = (normal / 20).max(10); assert!( @@ -543,7 +543,7 @@ mod strings_and_inline { } } assert!(total > 100, "expected many string hits, got {total}"); - // We intentionally don't assert non_utf8 > 0 — a stdlib hello binary + // We intentionally don't assert non_utf8 > 0 - a stdlib hello binary // may not contain any. The point is the API doesn't drop them. let _ = (utf8_ok, non_utf8); } @@ -776,7 +776,7 @@ mod buildinfo { } // =========================================================================== -// `moduledata` surfaces — segments, GC pointer maps, itabs, coverage, init +// `moduledata` surfaces - segments, GC pointer maps, itabs, coverage, init // =========================================================================== mod moduledata { use super::*; @@ -831,7 +831,7 @@ mod moduledata { /// `-buildmode=pie` renames every read-only-relocatable Go section with a /// `.data.rel.ro` prefix. If the classifier does not strip it, a PIE binary - /// looks like it has no type sections at all — which for Go ≤1.26 is read + /// looks like it has no type sections at all - which for Go ≤1.26 is read /// as a Go 1.27 signal, and for 1.27 loses the positive `.go.type` signal. #[test] fn pie_relro_section_names_are_recognized() { @@ -856,7 +856,7 @@ mod moduledata { /// Go 1.26 stopped writing `textStart` into the pcHeader, leaving the slot /// zeroed. A moduledata-less 1.26+ binary must not report `runtime.text` as - /// address 0 — that turns every entry VA into a raw `entry_off` without any + /// address 0 - that turns every entry VA into a raw `entry_off` without any /// error surfacing. #[test] fn text_va_is_never_zero() { @@ -1168,7 +1168,7 @@ mod embed_and_fips { /// `embed.FS` recovery reads through the address-space view, which for /// wasm is the reconstructed linear-memory image rather than the file. - /// Narrowing that search with file-offset section ranges finds nothing — + /// Narrowing that search with file-offset section ranges finds nothing - /// silently, since a binary with no embeds legitimately returns empty. #[test] fn embedded_assets_recovered_from_wasm() { @@ -1310,7 +1310,7 @@ mod wasm { fn wasm_arch_resolves_correctly() { let data = load(BASIC_WASM); let bin = GoBinary::parse(&data).unwrap(); - // pclntab arch can't disambiguate (1,8) — reports X86_64. + // pclntab arch can't disambiguate (1,8) - reports X86_64. assert_eq!(bin.pclntab().unwrap().arch(), Arch::X86_64); // bin.arch() folds in the container format and resolves to Wasm. assert_eq!(bin.arch(), Arch::Wasm); @@ -1405,9 +1405,9 @@ mod wasm { mod types { use super::*; - /// `abi.MapType` has shipped three different shapes — bucket-based `hmap` + /// `abi.MapType` has shipped three different shapes - bucket-based `hmap` /// (≤1.23), Swiss tables (1.24-1.26), and Swiss with explicit key/elem - /// strides (1.27+) — and each descriptor's size feeds the position of its + /// strides (1.27+) - and each descriptor's size feeds the position of its /// trailing `UncommonType`. Sweep the corpus and assert every map /// descriptor is read with the layout its Go version actually emitted, with /// the version-specific fields populated and the others genuinely absent. @@ -1545,7 +1545,7 @@ mod types { /// Go 1.27 replaced the typelink table with a walk bounded by /// `moduledata.typedesclen`. Walking to `etypes` instead runs off the end /// of the typelink descriptors into non-typelink types, `type:.namedata.*` - /// blobs, and finally the inline itab array — surfacing unnamed junk types. + /// blobs, and finally the inline itab array - surfacing unnamed junk types. #[test] fn v5_type_walk_stops_at_typedesclen() { use gobin::structures::moduledata::ModuledataVersion; @@ -1578,7 +1578,7 @@ mod types { ); assert!( !t.name.is_empty(), - "{}: unnamed type at {:#x} — the walk over-ran its bound", + "{}: unnamed type at {:#x} - the walk over-ran its bound", f.path, t.descriptor_va ); @@ -1790,7 +1790,7 @@ mod types { if let Some(tag) = f.tag { tagged += 1; // Conventional `key:"value"` tags (e.g. json/yaml) appear in - // stdlib structs — at least one should decode cleanly. + // stdlib structs - at least one should decode cleanly. if tag.contains(":\"") { saw_keyvalue_tag = true; } @@ -1809,29 +1809,29 @@ mod types { // produced a garbage `mcount`; resolving those phantom methods followed // garbage `mtyp` offsets and exploded the transitive closure. Real Go types // have at most ~150 methods, and this fixture's legitimate closure is ~850 - // types — neither bound is tight, but each is far below the bug's output + // types - neither bound is tight, but each is far below the bug's output // (mcount up to 0xFFFF; ~12.8k inflated types). for t in &all { assert!( t.method_count < 500, - "absurd method_count {} on type {:?} — UncommonType mis-located?", + "absurd method_count {} on type {:?} - UncommonType mis-located?", t.method_count, t.name ); } assert!( all.len() < 3000, - "all_types inflated to {} — garbage method offsets being followed?", + "all_types inflated to {} - garbage method offsets being followed?", all.len() ); // The transitive walk reaches the binary's own tagged structs and decodes - // their field tags — covering json, yaml, multi-key, and `,omitempty`. + // their field tags - covering json, yaml, multi-key, and `,omitempty`. // // (This previously asserted ">20" tagged fields. That count was an artifact // of a func-type descriptor bug: the `UncommonType` for func types was // located short by the funcType struct's alignment padding, so a garbage // `mcount` was read and the phantom methods' garbage `mtyp` offsets were - // followed into ~12k unrelated descriptors — inflating `all_types` from the + // followed into ~12k unrelated descriptors - inflating `all_types` from the // ~850 genuinely-reachable types to ~12.8k and surfacing stdlib structs not // actually reachable from typelinks. With the layout fixed, the walk reaches // only legitimately-referenced types.) @@ -1925,7 +1925,7 @@ mod types { // A stdlib-linked binary has plenty of named struct fields and func params; // the leaf (concrete) referenced types should resolve to names. Method / // interface-method signature types are unnamed Go func types, so their - // `type_name` is expectedly None — covered by the doc contract, not here. + // `type_name` is expectedly None - covered by the doc contract, not here. assert!(resolved_field, "expected some struct field type names"); assert!(resolved_param, "expected some func param type names"); assert!( @@ -2120,7 +2120,7 @@ mod types { let data = load(BASIC_NORMAL); let bin = GoBinary::parse(&data).unwrap(); - // bin.types() is a streaming iterator — verify it composes with .take() / + // bin.types() is a streaming iterator - verify it composes with .take() / // .count() and yields a positive number of types. let count = bin.types().count(); assert!(count > 0, "binary should expose at least one type"); @@ -2146,7 +2146,7 @@ mod types { assert_eq!(format!("{}", PclntabVersion::Go118), "go118"); assert_eq!(format!("{}", PclntabVersion::Go120), "go120"); - // Arch — values match canonical GOARCH where they exist + // Arch - values match canonical GOARCH where they exist assert_eq!(format!("{}", Arch::X86), "386"); assert_eq!(format!("{}", Arch::X86_64), "amd64"); assert_eq!(format!("{}", Arch::Arm64), "arm64"); @@ -2277,7 +2277,7 @@ mod types { // // Fixtures are discovered by filename // (`_go__[_variant][.exe]`), so the whole corpus -// built by `tests/samples/build.sh` is swept automatically — adding a build +// built by `tests/samples/build.sh` is swept automatically - adding a build // line extends coverage with no test edits. The `basic` program is byte-for- // byte identical source on every Go release from 1.2 onward, so the assertions // below hold uniformly across the entire version range. @@ -2523,7 +2523,7 @@ mod matrix { files.insert(s.to_string()); } } - // Bounded and populated — the legacy-misparse bug produced millions. + // Bounded and populated - the legacy-misparse bug produced millions. assert!( (10..5000).contains(&packages.len()), "{}: package count {} out of range", @@ -2541,8 +2541,8 @@ mod matrix { } // No package name carries control bytes or non-ASCII garbage. (Some // compiler-generated symbols like `type..eq.struct {...}` legitimately - // contain spaces and braces, so only control/non-ASCII bytes — the - // signature of the original misparse — are rejected.) + // contain spaces and braces, so only control/non-ASCII bytes - the + // signature of the original misparse - are rejected.) for p in &packages { assert!( p.bytes().all(|b| b.is_ascii() && !b.is_ascii_control()), @@ -2907,9 +2907,9 @@ mod harness { } } // Every structural detail the extractor enumerates from typelinks. - // (Bare interface type descriptors are not enumerated here — they + // (Bare interface type descriptors are not enumerated here - they // are reachable only via itabs and pointer/elem references, asserted - // through `itab_pairs` below — so "interface" is not in this set.) + // through `itab_pairs` below - so "interface" is not in this set.) for k in ["struct", "map", "chan", "slice", "array", "pointer", "func"] { assert!(kinds.contains(k), "{path}: missing TypeDetail kind {k}"); } @@ -2923,8 +2923,8 @@ mod harness { ); } // NOTE: named tagged structs (Record) are enumerated only in their - // `*main.Record` pointer form; the underlying struct descriptor — - // with field names, tags, and the embedded flag — is referenced via + // `*main.Record` pointer form; the underlying struct descriptor - + // with field names, tags, and the embedded flag - is referenced via // `elem_va` but not yielded by `bin.types()`. Field-tag/embedded // extraction is therefore not asserted here (see TODO C-06). assert!(saw_variadic_func, "{path}: variadic func type not surfaced"); diff --git a/tests/samples/README.md b/tests/samples/README.md index 802ec66..fb7b155 100644 --- a/tests/samples/README.md +++ b/tests/samples/README.md @@ -41,7 +41,7 @@ detect (one `basic_go_linux_amd64` fixture per row, plus the variants): | `go127` | Go120 | V5 | inline itabs, no typelinks, `.go.type`/`.go.func` sections | The corpus separately pins the `abi.MapType` layout, which has changed **six** -times and is not tied to the moduledata version — so each era needs its own +times and is not tied to the moduledata version - so each era needs its own fixture. The integration test `types::map_descriptors_use_the_layout_of_their_go_version` sweeps the whole corpus and additionally asserts that all six eras are present, so this coverage cannot be lost silently: @@ -61,25 +61,25 @@ normalizes them; `go120` (`hashMightPanic` on an interface-keyed map) and ## Programs (sources under `src/`) -- `src/basic/` — the primary fixture, **byte-for-byte source-identical across +- `src/basic/` - the primary fixture, **byte-for-byte source-identical across every Go release from 1.2**: `main.main`, `main.worker`, `main.(*TestStruct).DoSomething` (both `//go:noinline` so the symbols survive every version), interfaces with multiple implementors (itabs), a goroutine + channel, `defer`, a closure, struct tags, `fmt`/`reflect`/`strings`. Also the source for the `_cover`, `_fips`, and wasm fixtures. -- `src/types/` — the full type-descriptor zoo: every `TypeDetail` kind +- `src/types/` - the full type-descriptor zoo: every `TypeDetail` kind (struct/map/chan-all-directions/slice/array/pointer/func-variadic) funnelled through `reflect.TypeOf` so each descriptor is reachable. -- `src/generics/` — Go generics (≥1.18): generic functions and types +- `src/generics/` - Go generics (≥1.18): generic functions and types instantiated at multiple concrete types (`//go:noinline`), so the shape-stenciled instantiations (`main.Sum[go.shape.int]`, `main.(*Stack[…]).Push`, `main.Pair[string,int]`) are emitted. -- `src/cgo/` — a CGO build (`CGO_ENABLED=1`, needs host `gcc`), exercising the +- `src/cgo/` - a CGO build (`CGO_ENABLED=1`, needs host `gcc`), exercising the `_cgo_*` / `_Cfunc_*` shim symbols and the `CGO_ENABLED` build setting. -- `src/embed/` — `//go:embed assets/*` (multi-file `embed.FS` with a nested dir +- `src/embed/` - `//go:embed assets/*` (multi-file `embed.FS` with a nested dir and a binary blob) plus a single-file `//go:embed` string. -- `src/minimal/` — smallest useful program (no `fmt`); near-empty metadata. -- `src/plugin/` — a `-buildmode=plugin` Go plugin. Built natively on darwin +- `src/minimal/` - smallest useful program (no `fmt`); near-empty metadata. +- `src/plugin/` - a `-buildmode=plugin` Go plugin. Built natively on darwin (CGO) it is a Mach-O dylib using **chained fixups** (`plugin_go126_darwin_arm64.so`); see below. @@ -88,17 +88,17 @@ normalizes them; `go120` (`hashMightPanic` on an interface-keyed map) and Two fixtures exist to cover inputs a triage pipeline actually sees, where the parser cannot fall back on the usual markers: -- **`basic_go127_linux_amd64_pie`** — `-buildmode=pie` prefixes every +- **`basic_go127_linux_amd64_pie`** - `-buildmode=pie` prefixes every read-only-relocatable Go section with `.data.rel.ro` (`.data.rel.ro.go.type`, and pre-1.27 `.data.rel.ro.typelink`). Without the prefix stripping in `formats::classify_section`, a PIE binary looks like it has no type sections at all. -- **`embed_go127_wasip1_wasm`** — the only fixture where the address-space +- **`embed_go127_wasip1_wasm`** - the only fixture where the address-space view is not the file. `embed.FS` recovery searches the reconstructed linear-memory image, so any attempt to narrow that search with file-offset - section ranges finds nothing — silently, because a binary with no embeds + section ranges finds nothing - silently, because a binary with no embeds legitimately returns an empty list. -- **`basic_go124_windows_amd64_noversion.exe`** — a byte-identical copy of +- **`basic_go124_windows_amd64_noversion.exe`** - a byte-identical copy of `basic_go124_windows_amd64.exe` with every `go1.24` literal zeroed, the way an obfuscator leaves one. PE never carries a `.typelink` section at any Go version, so with no version string the only layout hint points at the Go 1.27 @@ -110,7 +110,7 @@ parser cannot fall back on the usual markers: `build.sh` is the single, self-contained builder. It runs on a **linux/amd64 host** (e.g. `ssh dev-linux`) and downloads each Go toolchain from go.dev on -demand — no system Go, no container engine, no `golang.org/dl` helpers. A linux +demand - no system Go, no container engine, no `golang.org/dl` helpers. A linux host is required because the pre-1.16 toolchains have no darwin/arm64 build (and crash under qemu user-emulation), while every modern format cross-compiles cleanly from linux. @@ -131,8 +131,8 @@ One fixture needs a host the script cannot assume: ``` Embedded source paths are trimmed so no build-host paths leak in: Go ≥1.13 uses -the `-trimpath` build flag, and Go 1.4–1.12 use `-gcflags=-trimpath=` (the +the `-trimpath` build flag, and Go 1.4-1.12 use `-gcflags=-trimpath=` (the fixture programs are pure Go), both reducing the path to `main.go` / `command-line-arguments/main.go` / `./main.go`. Only **Go 1.2** (`go12`) embeds -the build-cache path — its `6g` compiler predates path trimming entirely. +the build-cache path - its `6g` compiler predates path trimming entirely. Tests match source files by basename, so the exact form is irrelevant. diff --git a/tests/samples/build.sh b/tests/samples/build.sh index bf192d4..8e1b862 100755 --- a/tests/samples/build.sh +++ b/tests/samples/build.sh @@ -4,7 +4,7 @@ # # ONE script for the whole corpus. It runs on a linux/amd64 host and is fully # self-contained: it downloads each released Go toolchain from go.dev on demand -# — no system Go, no container engine, no gotip bootstrap. A linux host is used +# - no system Go, no container engine, no gotip bootstrap. A linux host is used # because the pre-1.16 toolchains # have no darwin/arm64 build (and crash under qemu user-emulation), while every # modern format (Mach-O / PE / Wasm) cross-compiles cleanly from linux. The cgo @@ -151,7 +151,7 @@ build() { # accepts `-gcflags=-trimpath=`, which strips the build dir so embedded # source paths reduce to "main.go" (the fixture programs are pure Go, so no # `-asmflags` is needed). Only Go 1.2's `6g` has neither and embeds the - # (fixed, build-cache) path — a single fixture, by necessity. + # (fixed, build-cache) path - a single fixture, by necessity. local flags=() if [[ "$m" -ge 13 ]]; then flags+=("-trimpath") @@ -167,7 +167,7 @@ build() { done # Copy the whole program dir (so go:embed assets travel with main.go), but - # drop go.mod — every build runs in GOPATH mode (GO111MODULE=off) and a stray + # drop go.mod - every build runs in GOPATH mode (GO111MODULE=off) and a stray # module file would only confuse the older toolchains. rm -rf "$bdir"; mkdir -p "$bdir" cp -r "$SRC/$prog/." "$bdir/"; rm -f "$bdir/go.mod" @@ -232,7 +232,7 @@ for entry in "${VERSIONS[@]}"; do # `.data.rel.ro.typelink`). Without a PIE fixture the section # classifier's prefix stripping is untested, and an unprefixed # lookup makes a PIE binary look like it has no typelink section at - # all — which the moduledata arbitration reads as a Go 1.27 signal. + # all - which the moduledata arbitration reads as a Go 1.27 signal. build "$tag" "$ver" basic linux amd64 "_pie" -buildmode=pie build "$tag" "$ver" minimal linux amd64 "" build "$tag" "$ver" embed linux amd64 "" @@ -252,7 +252,7 @@ done # A Go binary with every occurrence of its version string overwritten, the way # an obfuscator (garble) or a repacker leaves one. With no version to key on, -# the moduledata parser has to pick its layout from structural evidence alone — +# the moduledata parser has to pick its layout from structural evidence alone - # and on PE, which never carries a `.typelink` section, the only hint points at # the Go 1.27 layout. Guessing wrong there silently empties the types, itabs, # init tasks and inline tree, so this fixture pins the arbitration.