Skip to content

Run the 6502 on a cycle-exact bus and emulate the C64's chips cycle by cycle - #357

Merged
highbyte merged 61 commits into
masterfrom
feature/cpu-cycle-engine
Sep 28, 2026
Merged

highbyte merged 61 commits into
masterfrom
feature/cpu-cycle-engine

Conversation

@highbyte

@highbyte highbyte commented Sep 28, 2026 •

Copy link
Copy Markdown
Owner

Replaces the C64's per-instruction timing with a cycle-exact model: every CPU cycle is a real bus access, and the VIC-II, CIAs and SID see each read and write at the cycle it happens. Merged here from 51 reviewed PRs (#305–#358).

CPU

  • Every CPU cycle is a real bus access, including dummy reads and writes, verified against the SingleStepTests corpus.
  • Interrupts are sampled at the cycle the 6502 polls them. An IRQ released in the poll cycle is still taken, and an NMI can hijack an IRQ or BRK entry.
  • Added the undocumented opcodes LXA, ANE, SHA, SHX, SHY and TAS (including page-crossing and RDY behaviour). ARR now runs in the stable profile.
  • Undriven 6510 port bits keep the charge they were last driven to.

VIC-II

  • The VIC-II stalls the CPU on bad lines and during sprite DMA.
  • Counters, the border flip-flops, DEN, display/idle state and the sprite counters are followed cycle by cycle.
  • A new default pixel generator follows the graphics sequencer pixel by pixel, so mid-line mode, scroll, colour and pointer changes land at the correct pixel.
  • Sprite X compare, priority, collisions, Y-expansion and crunching behave as the chip does.
  • The previous generator is kept as the faster "legacy" option.

CIA and SID

  • The 6526 timer pipeline, interrupt control register timing and interrupt output are modelled.
  • The VIC bank comes from CIA 2's port pins and changes on the cycle after the write.
  • SID register writes and OSC3/ENV3 reads happen on their exact cycle, and reads return the data-bus latch.

Tools and docs

  • A VICE test-program harness (tools/vice-testprogs) runs VICE's VICII, CIA, CPU and interrupt test programs against their reference pictures and results.
  • The debug adapter and monitor show the VIC-II raster position and gain run until <condition> and run raster <line> [cycle].
  • New samples were added, along with new Download & Run entries (Demus Interruptus, Krestage 3, Chars Sucks, Robot - Not Human).
  • A new page, docs/systems/c64/accuracy.md, describes what is emulated and its known limitations.

Results

  • VICE test programs: CPU, interrupts and Lorenz 413/437; VICII 226/274; CIA 31/104. Most of the CIA failures come from features that aren't modelled, such as TOD, SDR, CNT and cascade; accuracy.md lists them.
  • Performance, measured on Giana Sisters (render and audio, Apple M1): a frame takes 1.32–1.34 ms against 1.03 ms on master (+28–29%). About 0.07 ms is the cycle-exact CPU, bus and chips, spread over many small costs; the rest is the per-pixel sequencer. The legacy pixel generator keeps the increase at 6–7%. Details in benchmarks/Highbyte.DotNet6502.Benchmarks/RESULTS.md.
  • The published AOT Avalonia Browser build renders in about 7–8 ms per frame on an M1.

…#305)

First increment on the `feature/cpu-cycle-engine` integration branch.
Nothing here touches `src/`.

Adds `labs/CycleEnginePrototype/`: a self-contained lab that compared
candidate CPU execution engines for cycle-level timing against the
production executor on a representative instruction slice
(immediate/absolute/indexed reads, indexed store with dummy read, NMOS
and CMOS RMW, branches, NOP, PHA, JSR, RTS, RTI, IRQ/NMI entry), one bus
access per cycle, driven through a device stub with a PAL VIC-II
bad-line BA pattern and a CIA timer IRQ.

Outcome, recorded in the lab README with the measurements:

- The cycle-stamped atomic design is chosen: the instruction stays
atomic, every cycle is a real bus access carrying the cycle number, and
the bus inserts RDY stall cycles. Lazy device synchronization
(closed-form device advance at instruction boundaries and predicted
stall cycles) is the production policy; per-cycle synchronization is
kept as the test oracle. Cost: 1.03x the current executor with no
devices, 1.22–1.27x with per-cycle-exact devices, allocation-free.
- Two resumable candidates (a micro-operation table and a flattened
state machine) were built, measured at 1.15–1.24x without devices and
1.55–1.65x with per-cycle device ticking, and removed once the choice
was made. Their code is in this branch's history; their numbers stay in
the README as the evidence.

Tests (44) hold both device policies to hand-derived bus traces, the
production executor's final state and cycle totals, the stall rules
(reads stall while BA is low, writes never), and to each other over a
5,000-instruction run with bad lines and timer interrupts, device state
included. The lab projects are in the solution so the build-and-test job
covers them.
…epTests corpus (#306)

Second increment on the `feature/cpu-cycle-engine` integration branch.

Every instruction of every CPU model now performs exactly the bus
accesses the silicon performs, one per clock cycle, in hardware order.
Memory-mapped I/O therefore sees the same accesses a real device would,
at the same points in the sequence. This is the foundation the
cycle-level C64 timing work builds on: with one access per cycle, the
bus can carry the cycle number and decide stalls at the access.

**Handlers**
- Implied/accumulator: next-byte dummy read. Zero-page indexed and
(zp,X): un-indexed base read. Branches: next-opcode read when taken,
wrong-page read on a crossing. PHA/PHP/PHX/PHY: next-byte read.
PLA/PLP/PLX/PLY, RTS, RTI: next-byte and stack-top reads. JSR: stack-top
read and hardware fetch order.
- 65C02: operand re-reads on the extra cycles of indexed stores,
page-crossing reads and RMW, `JMP (ind)` and `JMP (abs,X)`, decimal-mode
ADC/SBC, and the defined NOPs.
- `CPU.BusCycles` counts accesses; `FetchWord`/`StoreWord` go through
the byte helpers.
- Fixed on the way: RRA, ISC and ARR ignored decimal mode.

**Verification**
- A pinned subset of the [SingleStepTests
65x02](https://github.com/SingleStepTests/65x02) corpus (MIT), 20
vectors per opcode for the `6502` and `wdc65c02` sets, vendored under
`tests/Highbyte.DotNet6502.Tests/Fixtures/SingleStepTests/` with the
upstream LICENSE, a manifest (commit + hashes) and
`tools/singlesteptests/extract.py` to regenerate it. Per opcode, final
state, the cycle-by-cycle bus trace, the reported cycle count and the
`BusCycles` advance must all agree. Every documented and
stable-undocumented NMOS byte and every NCR 65C02 byte match. Deviations
are listed in the test class with reasons: the seven unstable NMOS
opcodes no profile implements; the Rockwell bit instructions and WDC
`WAI`/`STP` that the NCR 65C02 executes as NOPs; the corpus's
fixed-address decimal extra cycle of immediate ADC/SBC; the `$5C` NOP
cycle count.
- `OneBusAccessPerCycleTests`: every defined byte of every model and
profile, from seeded random state, one access per reported cycle.
- CPU suite 1,404 passed; Systems suite 1,178 passed including the
real-ROM C64 boot; prototype lab 44 passed.

**Performance** (same machine as the recorded baseline): CPU-only
microbenchmarks 9–14% slower, proportional to the dummy reads in their
instruction mix (each added access is one delegate call through
`Memory`); integrated C64 frames within ShortRun noise. Details in
`RESULTS.md`.

**Docs**: the CPU library page gains a "Cycle and bus accuracy" section;
the developer docs describe the vector tests and how to regenerate the
fixtures.
## Summary

Reads of the SID's write-only registers no longer return 0. The chip has
no readable storage for them, but the last byte on its data bus lingers
and is returned until it decays. This branch models that latch, using
the CPU bus-cycle counter to time the decay.

- **Chip-wide latch.** Any SID register write, and any read of
POTX/POTY/OSC3/ENV3, loads the latch. A read of a write-only register
returns the latched byte while it is younger than about 7,400 cycles
(the measured 6581 lifetime), and 0 after that.
- **Read-modify-write on SID registers now works as on hardware.** A
fast loader's `DEC $D418` loading noise counts the master volume down
through all sixteen steps instead of resetting it to 15 on every byte.
In Giana Sisters this replaces the crackle of the previous increment
with the loading buzz a real 6581 produces.
- The 8580's much longer latch lifetime waits for a chip-model setting;
the decay constant is a single named value.

## Tests

New `SidRegisterReadbackTests`: latch read-back is chip-wide, decays
after its lifetime, `DEC $D418` counts down, reads of readable registers
load the latch, an untouched chip reads 0, and the filter registers read
back through the latch like every other write-only register. Systems
suite: 1,200 passed.

## Notes

Based on the integration branch after merging master's SID filter
register fix (#308), so the filter mapping and stability clamp come from
there.
#310)

## Summary

Extends the lazy device synchronization the SID uses to the VIC-II and
both CIAs. Each device keeps the CPU bus cycle it has been advanced to;
the C64 loop drives them from `CPU.BusCycles` instead of instruction
cycle deltas (identical totals: every cycle is one bus access, interrupt
entry included), and every VIC-II and CIA register mapping first catches
the device up to the cycle of the access.

Before, a register access saw the device as it was at the previous
instruction boundary.

## What changes for programs

- `$D012`/`$D011` reads return the raster line at the cycle of the read,
so the instruction's shape matters as on hardware (`LDA abs` reads on
its 4th cycle, `LDA (zp),Y` on its 5th).
- Raster-compare, `$D019` and `$D01A` writes act on their own cycle.
- CIA timer and interrupt-status reads see the count at the read cycle,
in both timer modes; a control write starts counting on its cycle.
- A device interrupt that becomes due during an instruction is serviced
right after that instruction, instead of one instruction later.
- In `UpdateEachRasterLine` mode the CIAs advance to the bus cycle of
each raster line change (and at every CIA access) instead of by a fixed
line length, so timer interrupts no longer arrive up to a line late when
the program touches the CIA.

Rendering is unchanged: the rasterizer still samples registers once per
raster line.

## Details

- The CPU reset's vector fetch is excluded from device time: devices are
realigned with the bus counter after a reset and after a snapshot
restore. No snapshot format change.
- `Vic2.AdvanceRaster` and `CiaBase.ProcessTimers` keep their delta
semantics for tests and tooling; the C64 uses `CatchUpTo`.
- Not modelled: the 6526's start-up pipeline delay (a timer counts from
the write's own cycle).

## Tests

New `C64DeviceAccessTimingTests`: raster read at the exact cycle (line
change on vs. after the read cycle), read cycle depends on instruction
shape, CIA timer reads at the read cycle in both timer modes, control
write starting the timer, raster interrupt due mid-instruction serviced
right after it, devices advancing by exactly the instruction's cycles
after a snapshot restore. Systems suite: 1,208 passed, including the
real-ROM boot tests.

## Performance

Measured on Apple M1 (MacBook Air), .NET 10.0.7, against the integration
branch (`C64ExecuteFrameBenchmark`, `C64ExecuteInstructionBenchmark`),
no allocations added:

| Benchmark | baseline | this branch | ratio |
|---|--:|--:|--:|
| Frame, CoreOnly | 221.6 us | 230.0 us | 1.04 |
| Frame, RenderOnly | 408.9 us | 403.2 us | 0.99 |
| Frame, AudioOnly (sprites) | 376.8 us | 375.3 us | 1.00 |
| Frame, RenderAndAudio | 627.3 us | 596.3 us | 0.95 |
| Frame, RenderAndAudio (sprites) | 620.6 us | 604.8 us | 0.97 |
| 1000 instructions, CoreOnly | 25.1 us | 25.0 us | 1.00 |
| 1000 instructions, RenderOnly | 46.4 us | 46.6 us | 1.00 |
| 1000 instructions, AudioOnly | 44.9 us | 43.0 us | 0.96 |
| 1000 instructions, RenderAndAudio | 72.7 us | 71.7 us | 0.99 |

Within run-to-run noise on this laptop. Cost per VIC-II or CIA register
access is one bus-cycle comparison and, when cycles have passed, the
existing advance code.
## Summary

The CPU took any active IRQ or NMI at the next instruction boundary. The
6502 polls its interrupt inputs at the end of an instruction's
**second-to-last** cycle (a taken branch that does not cross a page: at
the end of its first cycle), so a line that goes active during the last
cycle is only seen after the following instruction. This PR models that.

- **Assertion cycles.** `CPUInterrupts` records the bus cycle on which
the IRQ line went active (first source to pull it low) and on which the
NMI edge was latched, via new overloads taking a cycle. Sources set
without a cycle keep the previous next-boundary behaviour, so other
systems and cartridges are unaffected.
- **Poll point.** The CPU records each instruction's poll point from
`CPU.BusCycles` and takes an interrupt at a boundary only if the
assertion cycle is at or before it. After an interrupt entry the poll is
re-armed at its end.
- **CLI, SEI, PLP vs RTI.** The first three change the I flag after the
poll, so their boundary decision uses the flag from before the
instruction (CLI's effect is one instruction late, an IRQ pending at SEI
is still taken). RTI changes it in time. Which opcodes are branches or
flag-changers comes from the model's opcode descriptor table
(`Addressing == Relative`, new `ChangesInterruptDisableAfterPoll`
marker), not from opcode literals in the CPU.
- **C64 devices date their assertions**: the VIC-II to the cycle the
raster line began, CIA timers to the underflow cycle, a `$D01A` enable
to the write cycle. The C64 loop now catches the VIC-II up **before**
the pending-interrupt check, which closes the case the previous
increment left open: an interrupt due during an instruction with no I/O
access is serviced right after it instead of one instruction later.

Reference: nesdev "CPU interrupts" (agrees with 64doc) for the poll
point, the taken-branch case and the CLI/SEI/PLP/RTI behaviour.

## Not modelled

- NMI hijacking an interrupt sequence already in progress.
- The 6526's one-cycle delay between flag and IRQ output.
- Pre-existing: enabling a CIA mask bit while its flag is already set
does not assert the line.
- In the default `UpdateEachRasterLine` timer mode a CIA underflow
between line changes is still discovered at the next line change (that
mode's known laziness).

## Tests

- New `CPUInterruptSamplingCycleTests` (13): IRQ and NMI asserted on the
second-to-last vs the last cycle, the unstamped default,
first-source-wins, taken branch with and without page crossing, CLI,
SEI, PLP, RTI.
- C64 `C64DeviceAccessTimingTests`: raster interrupt due on cycle 3 vs
cycle 4 of a 4-cycle instruction, with and without an I/O access; CIA
timer interrupt dated to its underflow cycle.
- Solution: 2,800 tests passed (6 projects).

## Performance

Apple M1 (MacBook Air), .NET 10.0.7, same machine for both runs, no
allocations added:

| Benchmark | integration branch | this branch | ratio |
|---|--:|--:|--:|
| CPU, 1000 NOPs, `ExecuteOneInstructionMinimal` | 15.16 us | 15.61 us |
1.03 |
| CPU, 1000 NOPs, `Execute` | 15.89 us | 16.36 us | 1.03 |
| C64, 1000 instructions, core only | 25.15 us | 24.21 us | 0.96 |
| C64, 1000 instructions, render and audio | 71.78 us | 71.55 us | 1.00
|

About half a nanosecond per instruction on the bare CPU loop (the
poll-point bookkeeping), within noise at C64 level.
…#312)

Follow-up to #311. The binding helper for implied instructions had grown
to eight parameters (Sonar S107, Major). The three descriptors that
change the I flag after the interrupt poll (CLI, SEI, PLP) are now
re-issued with the marker set after they are bound, and the extra
parameter is gone from the `Implied` and `Bespoke` helpers. No behaviour
change; CPU suite 1,428 passed.
)

## Summary

The CIAs were advanced from two places depending on `TimerMode`: the C64
loop after each instruction, or `Vic2.AdvanceRaster` on a raster line
change (the default, where a timer interrupt could arrive up to 63
cycles late). This PR makes the C64 instruction loop the single stepping
site, makes exact per-instruction stepping cheap, and retires the mode.

**Commit 1 — one stepping site, measurement.** The C64 loop steps the
CIAs in both modes (raster-line mode: when the line changed since the
last step). Timer storage goes from a dictionary to two fields. The
frame and instruction benchmarks got a `TimerMode` parameter, and the
benchmark machine now seeds CIA1 timer A the way the KERNAL does (latch
`$4025`, counting, interrupt masked) so the numbers reflect a real
machine. Measured cost of the old per-instruction mode over raster-line
mode with a counting timer: +27% core-only, +5–6% with render and audio.

**Commit 2 — deadline timers, mode retired.** A counting timer is held
as the bus cycle at which it underflows (counter derived on read), and
each CIA keeps the earliest such cycle, so the per-instruction catch-up
is two comparisons and a store in an inlined method. `TimerMode` is
removed from config, `C64`, the snapshot module's check, benchmarks and
tests. The CIAs are exact at every instruction boundary and every
register access in the shipped configuration.

## Performance

Apple M1, .NET 10.0.7, same session, CIA1 timer A counting:

| Frame scenario | raster-line mode (before) | old per-instruction |
exact (now) | Δ |
|---|--:|--:|--:|--:|
| CoreOnly / None | 224.8 µs | 285.0 µs | 226.9 µs | +1% |
| CoreOnly / MixedVisibleSprites | 233.1 µs | 267.2 µs | 235.0 µs | +1%
|
| RenderOnly / None | 394.9 µs | 444.0 µs | 408.5 µs | +3% |
| AudioOnly / None | 370.6 µs | 408.9 µs | 376.8 µs | +2% |
| RenderAndAudio / None | 588.2 µs | 622.8 µs | 593.3 µs | +1% |
| RenderAndAudio / MixedVisibleSprites | 603.3 µs | 634.2 µs | 607.2 µs
| +1% |
| `ProcessAllCiaTimers_Stopped` (both CIAs, per call) | — | 7.8 ns | 2.2
ns | −72% |
| `ProcessAllCiaTimers_Running_NoUnderflow` | — | 8.2 ns | 2.1 ns | −74%
|

A first deadline version that kept per-timer positions and called into
each timer every instruction measured no better than the old
per-instruction mode (+9–22%): the cost was six small method calls per
instruction, not arithmetic. Recorded in the RESULTS.md history.

## Behaviour changes

- A timer interrupt is now raised on the instruction boundary after its
underflow cycle (and dated to that cycle for the CPU's sampling rule)
instead of up to a raster line later.
- A continuous timer with latch N has a period of N+1 cycles as on the
6526. The old counter treated a counted-down 0 as a 65,536-cycle wrap,
which effectively stopped period-2 timers after their first underflow.
`C64CiaDummyReadTests` relied on that and now arms a one-shot timer.

## Snapshots

No format or version change. `c64-core` keeps the int slot that held the
mode (written 0, ignored on read; old snapshots load without a warning).
`c64-cia` still stores latch/control/current/running; the counter is
derived at capture and the deadline rebuilt at restore. Round-trip tests
pass.

## Tests

New `CiaTimerCountingTests`: period of latch+1, one-shot stop, live and
frozen counter reads, force load. Solution: 2,803 passed (6 projects).
## Summary

The VIC-II takes the bus from the CPU on hardware: 40 cycles on every
bad line (BA low from cycle 12, video matrix fetches in cycles 15–54)
and two cycles per sprite with DMA on, BA low three cycles ahead. The
6510 freezes on its first read while BA is low and resumes when it goes
high; writes never stall. Nothing modelled this, so the CPU ran 10–15%
faster than the machine and raster effects that depend on bad-line
timing landed in the wrong place. This PR models it.

- **CPU** (`IBusStallSource`, `CPU.BusStallSource`): before a read the
CPU asks how long the bus is busy (one comparison per read against a
next-check cycle the source names), adds the wait to `BusCycles` and to
the instruction's cycles, and performs the read at the release cycle.
Interrupt-entry reads can be stalled too. Writes are untouched.
- **VIC-II** (`Vic2BusStalls`): the BA-low windows of a raster line come
from the read's raster position and the current registers — the bad-line
window when DEN is set and `(line & 7) == YSCROLL` within `$30`–`$F7`,
and per sprite with DMA on a window around its pointer fetch (sprites
0–2 at the end of the line, 3–7 at the start of the next, so the
previous line's sprites spill into the first cycles); overlapping
windows merge so adjacent sprites keep BA low.
- **Sprite DMA is per-sprite state** on `Vic2` (`SpriteDmaMask`): it
switches on when an enabled sprite's Y equals the raster line at the
compare and runs for the sprite's 21 rows (42 when Y-expanded) whatever
the registers do meanwhile, as on the 6569. A Y written to a line the
raster has already passed costs nothing until the raster comes round
again — multiplexers rely on this (Commando parks its sprites at Y=30
from a handler running at line 30). Disabling a sprite mid-run does not
stop its fetch.
- **`AdvanceRaster` processes every line an advance crosses**, not only
the one it lands on. A read stalled by a bad line with sprites can span
a whole raster line, and a skipped line used to lose its raster compare
and per-line sprite snapshot.
- **Frame wrap** keeps the cycles an instruction ran past the frame end
(previously dropped; with stalls that could be most of a line) and
happens before the line is derived, so line 0 and a line-0 raster
interrupt are current in the same advance.

Register writes, the frame wrap and a resync invalidate the CPU's
next-check cycle.

## Field tests

- **Commando (in-game, PAL and NTSC)**: an earlier version of the sprite
model ("DMA on while the line is within 21 of Y", live registers)
charged eight sprites' fetches to lines the hardware never spends them
on; the game's interrupt chain lost a frame every other frame (sprites
vanished, status bar mangled, half speed). Reproduced headlessly from a
snapshot (298 of 300 frames alternating), bisected to sprite stalls,
traced to the late arming of the line-50 interrupt, fixed by the
per-sprite DMA state above. Now 0 alternating frames, same as the
integration branch; confirmed on the desktop.
- **Giana Sisters title, PAL**: clean before and after.
- **Giana Sisters title, NTSC**: one-frame white cells in the big
scroller every eight frames. This is the PAL title's raster-timed column
update overrunning the shorter NTSC frame once the CPU loses its
bad-line cycles; the old, faster CPU hid it. Hardware overruns here too,
but shows the old column for a row because it latches the video matrix
per character row during the bad line; our rasterizer reads
screen/colour RAM live per line, so the same overrun shows as a mixed
cell. The per-row matrix latch is the natural next render-side step. The
app's default variant is NTSC; most C64 software is PAL.

## Not modelled

The DEN latch during line `$30` and the display-off idle state (live DEN
is used).

## Performance

Apple M1 (MacBook Air), .NET 10.0.7, same session, no allocations added.
The benchmark machine has the display off, so the frame columns measure
overhead only; the last column is the same branch with `$D011 = $1B`.

| Benchmark | Baseline | After | Δ | After, display on |
|---|--:|--:|--:|--:|
| `CPU_Run_1000Instructions` | 15.13 µs | 16.20 µs | +7% | — |
| `CPU_Execute_NoSubscribers_1000Instructions` | 16.66 µs | 17.00 µs |
+2% | — |
| C64 frame `CoreOnly` / `None` | 231.0 µs | 231.2 µs | 0% | 218.3 µs |
| C64 frame `RenderOnly` / `None` | 399.9 µs | 401.2 µs | 0% | 388.1 µs
|
| C64 frame `RenderAndAudio` / `None` | 590.0 µs | 604.4 µs | +2% |
585.3 µs |
| C64 frame `RenderAndAudio` / `MixedVisibleSprites` | 598.9 µs | 609.3
µs | +2% | 594.5 µs |

The per-read comparison is the whole overhead (a NOP is two reads, hence
the CPU loop rows). With the display on the stalls take about 6% of the
CPU's frame, the hardware's share for 25 bad lines. Recorded in the
RESULTS.md history.

## Tests

`CPUBusStallTests` (6, scripted source): stall adds to bus cycles and
the instruction, shifts later accesses, writes not consulted, next-check
honoured, `RequestBusStallCheck`, source removal. `Vic2BusStallTests`
(24): bad-line window edges, unstalled writes, no-bad-line cases,
YSCROLL change, sprite 0 / adjacent sprites / sprite 3 wrap into the
next line, 21/42-line DMA span, DMA switching on at the compare and off
after its rows, a late Y write doing nothing, disable mid-run not
stopping DMA, a raster interrupt on a line jumped over by a stall, frame
boundary. Plus the wrap-remainder test in `C64DeviceAccessTimingTests`.
Solution: 2,834 passed (6 projects).
…le (#315)

## Summary

- The VIC-II rasterizer reads a character row's 40 screen codes and
colour RAM nibbles on the row's first line and reuses them for the
remaining seven lines, as the VIC-II does. Writes to screen or colour
RAM while the raster is inside a row become visible on the next frame
instead of mid-row. Bitmap mode uses the same latch for its colour byte.
A 40-column "fetched" mask validates the latch, so partial passes (row
starting off-screen, YSCROLL change mid-row) fall back to live reads
until a full line has been read.
- When a read is stalled by the VIC-II (`Vic2BusStalls`), the VIC-II and
the render provider are caught up through the stalled cycles before the
CPU continues, so a bad line's row fetch during the stall sees memory
before the stalled instruction's write, in hardware order.
- New sample `samples/Assembler/C64/Raster/row_latch.asm` (built `.prg`,
report and labels included). A raster interrupt busy-waits down the
screen and writes to rows the raster is inside: row 1 screen codes
cleared on line 3, row 2 colour RAM set red on line 3, row 3 a different
glyph on each of lines 1-7, row 4 the background colour register on line
3 as a non-latched control. All rows are restored in the bottom border.
Listed as "RowLatch" in the C64 menu's assembly examples in the Avalonia
and Blazor apps.
- Docs: `docs/systems/c64/libraries.md` describes the row latch and the
stall-time render catch-up.

## Verification

- New rasterizer test: a screen write three lines into a row leaves that
row unchanged and shows in the next row. Systems tests 1,241 and the
full solution 2,835 pass.
- Headless run of the sample on PAL and NTSC on this branch and on the
integration branch: with the latch rows 1-3 are pixel-identical across
all eight lines and row 4 splits at line 3; without it rows 1 and 2
split at line 3 and row 3 shows seven glyph stripes.
- Giana (NTSC) re-checked: the remaining one-frame white cells are the
game writing screen codes before colour RAM with the row fetch landing
between the two, which hardware shows too.
- Frame benchmarks back to back against the integration branch are
within noise (ratios 0.97-1.02 across core-only, render-only,
render-and-audio and sprite scenarios), measured on a MacBook Air Apple
M1, macOS, .NET 10. No RESULTS.md entry.
- Avalonia desktop and Blazor WASM apps build clean.
## Summary

- The VIC-II reports every register write made through the memory map
with the frame cycle it lands in and the register's canonical address.
The rasterizer journals those writes and applies the colour registers
(`$D020`-`$D024`) from the cycle after each write, instead of sampling
them once per raster line, so a colour change in the middle of a line
splits that line at the write's pixel position. Colour state is
resynchronised from register storage at the end of every frame, which
covers snapshot restores and writes that bypass the memory map. The
background layer's border and standard-text background are drawn as runs
that close where a colour write lands, so a line without colour writes
costs the same as before.
- The other display registers (mode, scroll, 38/24-column, memory setup)
are still sampled once per raster line.
- Two assembly samples, listed in the C64 menu's assembly examples in
the Avalonia and Blazor apps:
- **RasterColumns** (`raster_columns.asm`): two bands of vertical colour
columns in the top border showing how finely the colour can be changed
along a line. Four cycles gives 32-pixel columns with three preloaded
colours; six cycles gives 48-pixel columns with any colour. The 4-cycle
store is the floor, so 32 pixels is the finest a program can reach; the
VIC-II itself resolves 8.
- **ScreenColumns** (`screen_columns.asm`): the same columns across the
main screen with the display on. Only the bad line is lost, where the
VIC-II holds the bus across the line's whole visible width; each group
of 8 lines walks into that hold and is released on a fixed cycle, which
aligns it without counting the stolen cycles. Five of every eight lines
carry columns.

## Verification

- New tests: the write observer reports the cycle of the write and folds
mirrors to the canonical address; border and background splits land at
the block after the write on a top-border and a main-screen line; a
colour written on every line of a frame keeps landing at its own cycle
(the write journal is emptied as it is applied); colour registers set
outside the memory map are picked up at the end of the frame. Full
solution: 2,843 tests pass.
- Samples measured headlessly on PAL and NTSC: every column is exactly
32 or 48 pixels wide, positions are identical across groups, and
consecutive frames are pixel identical. Both samples were swept over
every program start cycle from 0 to 90 with no unstable phase, and 300
consecutive frames are identical.
- Frame benchmarks back to back against the base branch are within noise
(RenderOnly 443.9 to 428.1 us, RenderAndAudio 641.3 to 637.0 us, with
untouched scenarios moving 0.95 to 1.04 in the same runs), measured on a
MacBook Air Apple M1, macOS 26.6.2, .NET 10.0.7. No RESULTS.md entry.
- Avalonia desktop and Blazor WASM apps build clean.

## Known difference from other emulators

Our VIC-II model centres the 320-pixel drawable area within the raster
line, putting the first display pixel near cycle 11 where hardware opens
the display window around cycle 15. Cycle-timed colour changes therefore
land about 4 cycles further left than on hardware, a constant shift that
moves borders, sprites and text together and shows up only when
comparing against another emulator. Not addressed here.
…ine (#317)

## Summary

- The rasterizer placed the 320-pixel display window by centring it in
the raster line, which put its first pixel about four cycles earlier
than the VIC-II does. Every position derived from a cycle was shifted
sideways compared to hardware, while ordinary pictures looked the same
because everything moved together. It only showed when comparing
raster-timed effects against another emulator.
- Each chip variant now carries the offset of the display window into
the line (`Vic2ModelBase.DisplayWindowStartX`): 124 pixels on both the
6569 and the 6567R8, taken from the display window at X 24 and where X 0
falls relative to the line's first cycle. The two chips differ in where
their X counter wraps (after 504 on the 6569, after 511 on the 6567R8
whose 520-pixel line holds one X value for an extra 8 pixels later in
the line), which is why a naive `(24 - X at cycle 1) mod line length` is
right for PAL and 8 too many for NTSC; the derivation is in the source
comments.
- The visible frame is placed so the display window keeps its position
inside it, so static pictures are unchanged and only raster-timed
effects move. Sprites are anchored to the display window (X 24) and are
unaffected; the 38-column deltas already match hardware and are
unaffected.
- Adds `ColorChangePixelDelay`, the offset between a colour register
write and the pixel that first shows it, calibrated by eye against VICE:
-11 on PAL, 4 on NTSC, default 0 for unmeasured variants. These are
provisional: a pipeline delay cannot depend on the TV standard, so the
two values differing means some other model-dependent timing difference
is being absorbed there. The base-class comment states how to settle it
(compare something that does not depend on the bad line hold or the CIA
line clock, on both models, against VICE).

## Verification

- New tests pin the display window at 124 into the line on both models,
the visible-area placement derived from it, the unchanged border widths
inside the frame (41/42 PAL, 49/49 NTSC), and the cycle-to-pixel mapping
of a colour write against the chip figures. Existing colour split tests
derive their expected positions from the constants. Full solution: 2,851
tests pass.
- Rendered before and after headlessly: the plain BASIC screen is
byte-identical on both models; a sample mixing static text with one
raster-timed write differs on exactly the line with the write; the
column samples move by the predicted amounts and remain stable across
every program start cycle tested.
- Side-by-side against VICE on PAL and NTSC with the screen column
sample: the colour edges now coincide with VICE's on both models at the
calibrated values.
- Avalonia desktop and Blazor WASM apps build clean.

## Known open point

PAL at -11 and NTSC at 4 are 15 pixels apart, close to the two cycles by
which the line lengths differ. That points at a remaining
model-dependent timing difference against VICE, most plausibly in the
raster counter's timing or the sample's per-line path rather than the
bus hold. Being investigated separately; the source comments document
the procedure.
)

## Summary

The VIC-II's vertical state is now modelled per raster line instead of
being derived from the raster line and YSCROLL:

- **DEN (Display Enabled) latch**: DEN is sampled during raster line 48
and decides the frame's bad lines. Clearing it before that line switches
the display, and its bad lines, off for the whole frame; clearing it
later has no effect until the next frame. The bus stall source uses the
same condition.
- **Display / idle state**: a bad line starts a character row (RC = 0),
a row's eighth line drops the chip to idle state with VCBASE advanced by
40, VCBASE resets in line 0. The rasterizer draws rows from VCBASE/RC
and shows the byte at `$3FFF` (`$39FF` with ECM) in black in idle state.
- **Vertical border flip-flop**: set at the bottom compare line (251, or
247 with RSEL clear), reset at the top one (51/55) only if DEN is set.
While set, the whole line is border colour, sprites included. Missing
the bottom compare line by switching RSEL keeps the border open, and
sprites parked in the top and bottom borders are shown, on both the
per-line and the end-of-frame sprite path.
- A `$D011` write before a line's decision cycles (border at 16, bad
line at 14) re-decides that line.

Vertical fine scroll, FLD-style row stretching, a switched-off screen
and an opened border now come out as on hardware. The old
`_charGridYOffset`/`_scrollY` arithmetic and the invalid-band snap are
gone.

## Sample

`samples/Assembler/C64/Raster/idle_graphics.asm`, in the C64 menu of
both apps as **IdleGraphics**. One cycle-counted colour column loop runs
in every phase; the picture depends only on whether the frame has bad
lines:

1. DEN set: the text screen over columns that shear at every bad line.
2. DEN cleared on line 40 and set again inside line 48: the same, since
a write during the sample line counts.
3. DEN cleared through line 48: no bad lines, straight columns, and the
loop's per-line `$3FFF` write shows as a glyph in every cell (idle
graphics).
4. DEN cleared all frame: border only, columns edge to edge.
5. As 3 with 24 rows selected around line 251: the border never closes,
and eight sprites in the top and bottom borders appear.

The first phase waits for the space bar; the others follow at about two
seconds each.

## Verification

- 17 new tests: display state (bad lines, DEN decision, VC hold, border
compare lines, early YSCROLL writes), render (YSCROLL placement, idle
output, display off, border open below the screen, row resumed after an
invalid band), DEN latch bus stalls, sprites in an opened bottom border
on both paths, a sprite beginning above the visible area (PAL and NTSC,
both paths), a sprite running past the NTSC frame end.
- Full suite green (one wall-clock instrumentation timing test is flaky
under load and passes alone).
- Commando (two snapshots, D64 PAL/NTSC) and Giana Sisters (D64
PAL/NTSC) render byte-identically to the integration branch.
- The sample is frame-stable on PAL and NTSC at several start phases,
and both sprite paths render it pixel-identically.

## Docs

`docs/systems/c64/libraries.md` describes the DEN latch, display/idle
state and the border flip-flop instead of listing them as not modelled.
…hange delay against VICE (#319)

## Summary

Two related changes that close the colour change offset question.

**Samples start on the same cycle on every run.** The cycle-timed
samples (`raster_columns`, `screen_columns`, `idle_graphics`) took their
line-clock reference from the first arrival of the run, which depends on
the cycle the program was started on, so the columns landed on a
different cycle from launch to launch: two phases on PAL and three on
NTSC for `screen_columns`. That made screenshots incomparable between
emulators and had produced two different per-model calibrations. Now:

- The reference is taken from the raster line change itself: an 11-cycle
poll over 16 top-border lines (a length neither 63 nor 65 is a multiple
of, so every loop position comes up), keeping the reading from the
earliest exit; every reading is stored and reduced afterwards so the
drift is constant, and the reduction handles readings that straddle the
timer's reload.
- The timer difference is mapped to the delay slide entry through a
page-aligned table instead of masked arithmetic, which was wrong near
the reload; constant time, both period-shifted representations covered.
- `screen_columns`' walk-in arrival turned out to be a per-model
constant against that reference and is built into the table.
`idle_graphics` enables its sprites after the calibration, since their
DMA stalled its polls.

Verified with a start-phase sweep of 27 phases on both models: one first
colour write cycle per sample and model, and the same result with the
timer's phase shifted through its reload in test builds.

**One measured colour change delay.** With comparable screenshots, VICE
x64sc (CRT off) shows every colour edge within a pixel of this
emulator's on PAL (6569 and 8565) with -11, and 15 pixels left of this
emulator's on NTSC (6567R8) with the previous +4. So
`ColorChangePixelDelay` is -11 on both models, in the base class, with
the per-model overrides removed. Its doc comment explains the
measurement, what the value means (five pixels of pipeline plus about
two cycles of alignment between this emulator's cycle index and the
chip's pixel timeline, the latter not yet located), the candidate
explanations, and how to pin it down.

## Verification

- Start-phase sweep, 27 phases, PAL and NTSC: one first-write cycle per
sample and model.
- Calibration wrap: test builds shifting the timer start through
reference values 0, 1, 8, 11 give the same result.
- Frame stability on both models for all three samples; pictures
unchanged in composition.
- NTSC edge positions after the change: 9, 41, 73 and 265 pixels into
the display area against VICE's 8.5, 40.8, 72.6 and 264.7; PAL unchanged
and matching.
- 404 C64 tests pass; Sonar clean.

## Docs

`docs/systems/c64/libraries.md`: the colour delay is described as
measured and the same on both models.
…sample (#320)

## Summary

**Border unit.** The rasterizer now follows the VIC-II's main border
flip-flop pixel by pixel instead of a per-line 38/40 column layout:

- Set when the X coordinate reaches the right compare value (344 with 40
columns, 335 with 38); reset at the left compare (24 or 31) on a line
the vertical flip-flop leaves open. A pixel is border colour while the
flip-flop is set. The ordinary 40 and 38 column layouts are what those
rules give on an unchanged line.
- The column select bit follows the register write journal at the cycle
boundary after the write (no colour pipeline delay), so a program that
selects 38 columns in the one cycle between the two right compares
misses both and keeps the side borders open on that line and the left
border of the next, as on hardware.
- An opened border shows the sequencer's idle output (the byte at
`$3FFF` in black over the background colour) on the character grid.
Graphics, the background prefill and sprites are clipped per line to the
span where the flip-flop is clear, kept per frame row for both sprite
paths. Character columns are drawn one cycle later than before so a
compare in the same cycle is already known.
- Sprite X positions wrap at 512 as on the chip, with the line starting
at X 404 (PAL) or 412 (NTSC): a sprite at X 496 sits in the left border.
- The 6567R8's sprite DMA fetches are one cycle later than the 6569's
(sprite 0 pointer fetch in cycle 59, from VICE's cycle tables); the
stall model had scaled the PAL offsets by line length and was two cycles
late on NTSC.

**Sample.** `samples/Assembler/C64/Raster/side_border.asm`, in the C64
menu of both apps as **SideBorder**: a band of 21 lines with eight
X-expanded sprites across the whole frame width, the outer ones in the
side borders, over the background running edge to edge. The band avoids
bad lines with FLD and is timed by the sprite DMA hold itself, whose
release is on the same cycle every line (10 on PAL, 9 on NTSC), so only
its first line is entered from the line clock sync.

**Sample fix.** The previous increment's `SyncToLine` clobbered X in its
table lookup, which broke this sample's line counter and had changed a
colour in both column samples. Fixed in all four samples; their cycle
comments now carry measured numbers (the sync returns 50 cycles into the
line).

## Verification

- New tests: side border opening with a 38-column write in cycle 54 and
not in 53, on both sprite paths, with a sprite in the opened right
border and a wrapped-X sprite in the opened left border; NTSC sprite DMA
window one cycle later than PAL and the release cycle for all eight
sprites on both models.
- 411 C64 tests pass; full suite green apart from a wall-clock
instrumentation timing test that is flaky under load and passes alone.
- The sample's writes land on their cycles on every traced line and at
every start phase on both models; frames are stable.
- Commando (two snapshots, D64 PAL/NTSC) and Giana Sisters (D64
PAL/NTSC) render byte-identically to the integration branch.

## Docs

`docs/systems/c64/libraries.md` describes the border unit, the per-cycle
column select, the sprite X wrap and the NTSC sprite fetch timing.
…tical border latch and raster IRQ edge (#321)

## Summary

A console tool under `tools/vice-testprogs` runs the VIC-II test
programs from the [VICE
repository](https://sourceforge.net/p/vice-emu/code/HEAD/tree/testprogs/)
on the booted C64 (real ROMs), stops at their `$D7FF` exit write and
compares the frame with the reference pictures pixel by pixel, on both
models. It writes a reference/ours/diff picture per test and a
`results.md`. The programs and references are not in the repository; the
development doc says where to fetch them and how to run it.

Fixes the runner revealed, each with unit tests:

- **Border compare cycle.** The border unit evaluates each compare one
cycle after the cycle its X coordinate falls in, with the registers as
written before that cycle (VICE's cycle tables: the 38 column right
compare in cycle 56 counting from 1, the 40 column one in 57). The
rasterizer's compare cycles moved by one and columns are drawn two
cycles after the one they start in. border-250/251/252 now match.
- **Vertical border latch.** The compares are checked every cycle with
the registers as they are then: the bottom compare arms a latch the
flip-flop takes over at line start and at the left compare, the top
compare with DEN clears both at once. A `$D011` write in a line's last
cycle is first seen on the next line. All vborder/vborder2 variants now
match.
- **Raster interrupt edge.** The interrupt is raised when the comparison
goes from non-match to match, checked at line entry (a cycle later for
line 0) and in the cycle after a `$D011`/`$D012` write. Writing the
current line's number raises it at once; a program that moves the
compare to the next line in every line's last cycle gets no further
interrupt (rasterirq_hold). Previously every line start re-raised it.
- **Per-line foreground clear** in the rasterizer, so NTSC's wrapped
rows no longer keep the previous frame's pixels.

Samples: the side border sample makes its 38 column write with DEC,
whose two writes fall in cycles 55 and 56 while the sprite DMA hold
stops the CPU's reads from cycle 55 (the same technique as VICE's
hvborder test); its slide centre is 9 like the other samples. The four
cycle-timed samples enter their calibration's first poll on the line
before the first polled one: entered on that line, the first two polls
got out at once and the reference came out wrong on most start phases.

## Baseline

| Suite | PAL | NTSC | Remaining differences |
|---|---|---|---|
| dentest | 21/21 | 21/21 | |
| border | 15/21 (2 without reference) | 2/2 without reference |
bm-idle, bm-ysh, bm-ysh2, mcbm, hvborder1/2: idle-state graphics with
mid-line XSCROLL/mode changes |
| D011Test | 13/14 (13 without reference) | 0/1 | disable-bad: mid-line
bad line disable |
| rasterirq | 1/1 | | |
| fldscroll | 1/7 | | mid-line bad line / VSP |
| screenpos, spritedma, videomode, gfxfetch, colorsplit | mostly differ
| | per-cycle sequencer, sprite Y-expand |

## Verification

- 415 C64 tests pass (5 new). Full suite green.
- Commando (two snapshots, D64 PAL and NTSC) and Giana Sisters (D64 PAL
and NTSC) byte-identical to the integration branch.
- Side border sample verified on both models and two start phases; the
other three samples' first-write cycles unchanged against the
integration branch.
- Sonar branch analysis clean.
…er pixel generator (#322)

## Summary

The Vic2Rasterizer gets a new default pixel generator,
`Vic2RasterizerSequencerPixelGenerator`, that follows the chip's
graphics data sequencer pixel by pixel as the VIC-II article describes
it: each cycle's g-access feeds a two-stage pipeline, the shift register
loads at the XSCROLL pixel of the cycle after, the mode bits take effect
part way through a cycle (MCM at pixel 4, ECM and BMM at pixel 4 when
set and 6 when cleared, with the `$D023` flash for hires cells during an
MCM switch), and each pixel's colour source and priority come from the
article's mode tables. XSCROLL, the mode bits and `$D018` reach the
sequencer through the register write journal. Pixels are recorded as
colour codes and resolved into the layers when the line ends with the
background colour registers' values at each pixel, so `$D022`/`$D023`
mid-line writes are exact and bitmap clear pixels are background
priority. The points within a cycle at which register changes reach the
output are observed chip behaviour, with the VICE project's emulation
and test programs credited in the class comment as the reference; the
code is written from the article, not ported (the first draft was a
close port and was rewritten, verified pixel-identical across all 85
result pictures on both models).

The previous generator, `Vic2RasterizerUintPixelGenerator`, stays
unchanged as the legacy pixel generator: faster, selectable with
`C64Config.Vic2RasterizerPixelGeneratorType` (a checkbox in the Avalonia
and Blazor C64 settings, and `appsettings.json`), and receiving no new
features. Both implement `IVic2RasterizerPixelGenerator`. The Avalonia
config dialog now also copies the rasterizer options back on OK, which
it did not for the per-line sprites setting either.

## VICE test programs (PAL)

| Suite | Before | After |
|---|---|---|
| border | 15/21 | 21/21 (bm-idle, bm-ysh, bm-ysh2, mcbm, hvborder1/2
now match) |
| colorsplit | 0/1 | 1/1 |
| dentest, rasterirq | 21/21, 1/1 | unchanged |
| videomode | hundreds of pixels per test | 2-87 pixels per test, single
mode-switch edge pixels |
| gfxfetch | 448 | 448: the test changes character RAM in the cycles the
chip fetches it; the rasterizer runs after each instruction and sees the
change early (documented limit) |

## Performance

Apple M1, same session, integration branch as baseline. RenderOnly
frame: 360 -> 393 µs with the display off, 354 -> 394 µs with the
display on, after a constant-time path for cycles with nothing to show
and a branch-free steady-state path. Three further reductions (colour
codes reused per block, blank bytes filled directly, event-free lines
resolved through a lookup table) followed; their Release re-measurement
was inconclusive because of machine load and is noted as such in
RESULTS.md. In a Debug build the app's stats panel showed the render
provider's per-instruction time go from 2.27 ms to 1.77 ms per frame on
a BASIC screen, against 1.51 ms for the legacy generator.

## Verification

- 454 C64 tests (five new: mid-line XSCROLL, MCM at pixel 4 with the
`$D023` flash, `$D018` from the next g-access, bitmap pixel priority, a
multicolour pair held across the cycle boundary during an XSCROLL
change; plus the configured generator reaching the rasterizer). Full
suite green.
- Commando (two snapshots, D64 PAL and NTSC) and Giana Sisters (D64 PAL
and NTSC) byte-identical to the integration branch except two pixels in
Commando NTSC.
- The new config checkbox was verified by hand in the desktop app.
- Sonar branch analysis clean.

## Sample

`samples/Assembler/C64/Raster/line_splits.asm` ("LineSplits" in both
menus): three bands where every line writes a register 36 cycles in and
restores it after the right border compare, so each line is in the old
state on its left half and the new one on its right: XSCROLL from a
triangle wave, `$D018` to the lower case set, MCM on. Bad lines cannot
be split (the CPU is halted there) and show as a gap in the seam every
eighth line; that halt releases on a fixed cycle and keeps the loop in
step from row to row. With the legacy generator the same program shows
no split at all.
…RowStretch samples (#323)

## Summary

- The sequencer pixel generator keeps VC, VCBASE, RC and VMLI per cycle
as the VIC-II does. A bad line condition created in any cycle resets the
row counter and starts the c-accesses there, which gives the DMA delay,
linecrunch, doubled rows and FLD their real behaviour instead of a
per-line approximation. The vertical border still gates the graphics
data while the counters advance. The legacy generator is unchanged.
- The light pen line on CIA1 port B bit 4 latches LPX/LPY once per frame
and raises the VIC IRQ, so programs that sync a frame on it no longer
drift.
- Two new samples in the Avalonia and browser menus: **DmaDelay**
bounces the text screen sideways and vertically with a mid-line bad
line, **RowStretch** stretches a block-character logo by repeating
character rows through the row counter. Both wait for the space bar and
run one pass per press.
- **Fairlight Intro (Golden Collection)** added to the C64 Download &
Run lists (Avalonia, browser, terminal) with a new demos section in the
compatible programs docs.

## Validation

- 458 C64 tests pass, including new tests for the mid-line bad line
condition with its DMA delay and for the light pen latch (position, once
per frame, IRQ flag).
- VICE test programs on both models: fldscroll 7/7, dentest 42/42,
border 21/21, colorsplit, rasterirq and screenpos unchanged.
- Game snapshots (Commando, Giana Sisters, PAL and NTSC) render
byte-identical to the integration branch.
- Both samples verified per raster line on PAL and NTSC with a headless
probe; the picture after a pass is identical to the idle one.
- SonarCloud branch analysis clean at Major and above.
- Frame benchmark (MacBook Air Apple M1, .NET 10.0.203,
`tools/perf-compare.sh` against the integration branch, three runs):
RenderOnly without sprites 1.05 / 0.99 / 0.98, with sprites 1.03 / 1.25
/ 1.03 of baseline; CoreOnly swung from 0.95 to 1.07 between runs while
this branch hardly touches the CPU path. Within the day's measurement
noise; no consistent regression.
… from them (#325)

## Summary

- The VIC-II keeps each sprite's data counter, its base and the
expansion flip-flop as the chip does, with the checks in cycles 15, 16,
55, 56 and 58 applied as the raster advances. A change of the Y-expand
bit lands before or after the cycle 55 check according to its own cycle,
clearing the bit sets the flip-flop at once, and the DMA ends when the
counter base reaches 63. That gives the 23-line sprite of VICE's d017
tests and the sprite stretcher of the demos. The bus stall model
anticipates the compare that starts a sprite.
- With per-line sprites on, the sequencer pixel generator draws each
raster line's sprites from the bytes the chip fetched for it instead of
a band captured at the sprite's start, so pointer, data, position and
expansion changes mid-sprite show on the line they reach. The
frame-level sprite path and the legacy generator are unchanged.
- Performance: the sprite events skip all work while no sprite is
enabled or fetching and are applied behind one compare per raster
advance; the model's per-line constants are cached instead of fetched
through abstract properties on every advance; the generator counts its
line and cycle along instead of dividing per cycle.
- New sample **SpriteStretch** (a diamond stretched by a wave through
per-line `$D017` writes, in a bad-line-free band). Two sprite-only demos
in Download & Run, **For Your Sprites Only** and **Unfortunate
Coincidence**, with a `requiresPerLineSprites` flag that switches
per-line sprites on for a program.

## Validation

- 467 C64 tests, full suite 2935 passed, 9 skipped. New tests:
`Vic2SpriteDmaTests` (compare and display cycles, 21 and 42 rows, expand
bit cleared mid-sprite before and after the cycle 55 check, hold and
release of a row, redisplay by a rewritten Y) and a per-line render test
of the d017 case.
- VICE test programs: spritedma 4/4 on PAL and NTSC (were 0/4); 113
tests across ten suites with the same 18 known differences as before
(videomode residuals, gfxfetch, disable-bad).
- Commando (two snapshots, D64 PAL and NTSC) and Giana Sisters (D64 PAL
and NTSC) render byte-identical to the integration branch; one 60-frame
Commando run differs by 4 pixels where a sprite's last row is now
clipped by its own line's border span.
- Frame benchmark on the integration branch and this branch, alternating
(MacBook Air M1, .NET 10.0.203): without sprites level within the
run-to-run spread; with eight visible sprites +3 to +9%, the counters
and fetch state of fetching sprites per line.
- SonarCloud branch analysis clean at Major and above.
## Summary

The sequencer pixel generator runs its per-cycle pipeline over every
cycle of every line, so a frame that opens the vertical border, where
the whole line is output, is about twice the work of a text screen. On a
sprites-only demo that does so the sequencer rendered at three times the
legacy generator's cost. Four changes to the generator, none of which
change its output:

- **Identical-block cache**: a cycle whose fetched bytes, mode and
pipeline state (shift register, its matrix and colour, pixel value, pair
phase, XSCROLL load pixel) equal the previous block's copies its eight
foreground/background codes and takes that block's end state. Covers the
idle byte across an opened border, runs of identical cells, and blank
cells under any XSCROLL, which the aligned fast path never reached.
- The line's colour resolve copies such blocks' colours instead of
looking them up, via a per-line serial mark.
- Sprite rows are decoded into a buffer and written as runs through the
bulk pixel delegates instead of a delegate call per pixel.
- The fetch ring's slots are derived instead of taken as three modulos
per cycle.

## Validation

- Timing on this machine (MacBook Air M1, .NET 10.0.401, quiet, headless
probe running "For Your Sprites Only" frames 400-1000, medians of
alternating repeats): sequencer 1587 → 1172 µs per frame with per-line
sprites on, 1404 → 1025 off; legacy generator 826 / 725; CPU and VIC-II
alone 404.
- Frame benchmark RenderOnly rows, this branch vs the integration branch
alternating: 474-477 vs 476-489 without sprites, 491-494 vs 495-496 with
eight, i.e. level. An earlier version of the cache had cost the text
screen 7-10%; packing the compare into one key and using direct copies
removed that.
- Output identical: 113 VICE test programs across ten suites with the
same 18 known differences as before, 467 C64 tests, full suite 2935
passed, Commando and Giana Sisters byte-identical to the integration
branch, desktop and browser apps build clean.
- SonarCloud branch analysis clean at Major and above.
…teX sample (#327)

Compares every sprite's X register with the beam at every pixel, as the
VIC-II does, instead of once per line.

## What changes

- **X compare per pixel.** Writes to the sprite X and X-high registers
during a line are journalled with their cycle. When the line ends, each
sprite's output runs are derived from the X in force at each pixel: a
write counts from the fifth pixel of its cycle, so a sprite moved just
before the beam reaches it appears at the new X and one moved just after
stays at the old one.
- **The sprite's own fetch.** A sprite cannot start while its data is
being fetched (X 355-366 for sprite 0 on the 6569, two cycles on per
sprite). One still shifting when its fetch begins repeats its last pixel
for seven pixels and stops. The fetch loads the next row, which can
start again on the same line, so a sprite whose X lies beyond its fetch
shows each row a line higher. X values the beam never reaches (504-511
on the 6569) are never shown, and a row no match shifts out stays in the
register for the next line.
- **Collisions at line end.** Sprite-to-sprite collisions come from the
derived runs and are latched as the line ends, so a clearing read of
`$d01e` early in a line no longer wipes what the line's pixels produce.
Sprite-to-background collisions keep the start-of-line approximation for
now.
- **Sequencer generator.** Draws the runs (with a shown length and a
repeated last pixel) and no longer draws an enabled sprite the chip
never showed at its settled end-of-frame position.
- **Fix:** the sprite-to-background collision path read past the VIC
bank for a sprite line just below the text area (text row 25), an
`IndexOutOfRangeException` reachable from Demus Interruptus.
- **New sample `SpriteX`** (Avalonia and browser menus): all eight
sprites on a band in the opened lower border, a cycle-exact line loop
that writes two sprites a new X one pixel either side of their compare,
and the space bar stepping sprites 0 and 1 through the fetch-window
cases (whole, cut with the repeated pixel, not shown, a line higher)
with the collision register shown. Runs unchanged on real hardware for
comparison.

## Verification

- VICE `spritex` suite: 29/29 on PAL and 29/29 on NTSC (against the C64C
and 6567R8 tables in `testsuite.txt`); `demusinterruptus` check reads 3
on PAL.
- Harness: the tracked suites unchanged apart from spritex; the other
sprite suites compared against the integration branch only improved
(`spritecollisions/sprite-sprite`, `spritesplit/ss-xpos`, `spritey` now
match, nothing worse).
- Games A/B (Commando snapshots and disk, Giana Sisters, PAL and NTSC):
pixel-identical, including a previous 4-pixel difference that is gone.
- Tests: 479 C64 tests green, six new ones for the rules (write seen
from pixel 4, fetch window, frozen last pixel, next row on the same
line, X 504 on PAL vs NTSC, restart after the fetch); the per-line
collision tests now drive the raster.
- Benchmark (MacBook Air M1, two rounds): frame benchmark rows level
within noise; the eight-sprite demo's per-line path 5.5% (sequencer) and
7% (legacy generator with per-line collisions) slower, about 24 ns per
sprite-line, after making the line walk linear in the journal and
table-driving the collision masks. No cost with per-line sprites off.
- Demos: Demus Interruptus now runs through its first parts without
crashing; its FLI picture's lower half still needs the sprite crunch (a
separate increment). Krestage 3's emulator check gets bits 0, 2 and 3 of
`$d01e` from this change and still needs sprite splits for bit 1.
…s Interruptus (#328)

Models the VIC-II's sprite crunch and adds Demus Interruptus, whose FLI
picture depends on it, to Download & Run.

## What changes

- **Sprite data counter.** MC counts the line's three data fetches, so
it stands at MCBASE plus 3, and in cycle 16 a sprite whose expansion
flip-flop is set copies MC into MCBASE before the check that ends the
DMA at 63. This replaces the article's "+2 in cycle 15, +1 in cycle 16",
which described the same thing except in the crunch case; the cycle 15
sprite event is gone.
- **Sprite crunch.** A write that clears a sprite's Y-expand bit while
its flip-flop is clear, made in cycle 15, replaces MC with a bit-merge
of MCBASE and MC (odd bits where both are set, even bits where either
is) before cycle 16 copies it. The counter leaves its stride of three,
misses 63 and wraps, so the sprite's rows come out reordered and its DMA
runs on for up to 63 rows. The merge reproduces the table in VICE's
spritecrunch readme (0→1, 1→5, 3→7, 4/5/6→5).
- **Download & Run:** Demus Interruptus (Crest, 2001), a zipped disk
image, PAL, audio, per-line sprites required; the space bar moves
between parts, as its own scroller says.
- **Disk loader fix.** The Download & Run wildcard took the first
directory entry, which on this disk is a decorative deleted-type entry
(506 bytes of directory art), so RUN did nothing. It now picks the first
program file, as the drive's own wildcard does.
- Docs: the sprite paragraph in `libraries.md` and a row in
`compatible-programs.md`.

## Verification

- Demus Interruptus's picture is whole: its unrolled `$d018`/`$d011`
stream is paced by the bad-line stall plus crunched sprites that stall
the CPU on every line until the Y compare restarts them at line 192, and
the lower half was garbage before. Watched twelve minutes with space
presses through the sine bars, the picture and the side-border bar
parts: no exceptions, nothing garbled. Its first-file load verified
through the same content loader the menus call.
- Unit tests: the crunch in cycle 15 gives MC 7 where cycles 14 and 16
give 6 and 3; a crunched sprite's DMA runs to line 143 instead of 122;
the loader wildcard past a decorative entry. 483 C64 tests green.
- Games A/B (Commando snapshots and disk, Giana Sisters, PAL and NTSC):
identical.
- Harness (all suites, both models): unchanged except spriteenable3_ntsc
now matches, spriteenable4_ntsc and spritefetchbug closer,
spriteenable1/2_ntsc further off (not looked into), and sequencer-bug
384 → 8496 pixels. That last one is timing rather than the rule: its
stretched sprites need one crunch and run 82 lines, which our frames 10,
12 and 13 after start reproduce to the reference's last row; the harness
captures a frame like our frame 11, where both writes land one cycle
later and the sprites crunch twice. The interrupt entry phase in the
first frames differs from the reference's by a cycle, which belongs to
the CPU-level timing work. The spritecrunch suite itself stabilises on a
CIA timer and lands 35 cycles from where hardware has its writes in our
run, so it waits for the CIA timing milestone.
- Timing on the eight-sprite demo: level within noise (one sprite event
fewer per line).
… lists (#329)

## Summary

- The visible frame is now the chip's own: the lines and pixels outside
the vertical and horizontal blanking (VIC-II article, section 3.4). PAL
shows raster lines 16-299 from X 480, NTSC lines 41-12 from X 489.
Before, the frame was centred on the display window with a hardcoded
NTSC line offset.
- Each VIC-II model declares its first visible raster line and X; one
shared raster-to-screen-line conversion replaces the per-model overrides
(the old NTSC model no longer throws). `Vic2Screen` exposes the
asymmetric border sizes (PAL 35 above, 49 below, 48 left, 35 right; NTSC
10, 25, 47, 51); the symmetric values of the generic screen interface
stay for the text hosts.
- Pixels produced in the blanking, such as sprites parked at Y 0-15 on
PAL or an opened border in the line's last cycles, are no longer shown.
This removes two artefacts seen in Demus Interruptus: coloured "noise"
rows at the top and a white two-pixel strip at the right edge.
- Download & Run: the Terminal app's C64 list keeps only text-mode
programs (the demos it cannot display are removed); Demus Interruptus is
removed from every list; Smooth and Wonders (Singular, 2020) is added to
the Avalonia and browser lists, a 384x273 sprite hyperscreen with the
borders opened and sprites stretched by rewriting the Y-expand register
every line.
- Tests re-pinned to the new geometry; docs and comments updated.

## Verification

- Games A/B (Commando snapshot and D64, Giana Sisters D64, PAL and
NTSC): the new frame is pixel-identical to the old one shifted by the
window offset (PAL −7/+7, NTSC +2/+7), 0 pixels differ.
- VICE test-program harness, 274 tests: identical pass/fail statuses and
identical per-test pixel counts.
- Rendered border extents measured on the new frames match the figures
above on both models.
- Smooth and Wonders matches its csdb screenshot, ran 30000 frames
headless without an exception, needs per-line sprites (blank without
them) and also renders with the legacy generator.
- Whole solution builds; full test suite passes (the two wall-clock
timing tests pass when rerun alone).
…d Krestage 3 (#330)

## Summary

- The sprite priority (`$D01B`), multicolour (`$D01C`) and X-expand
(`$D01D`) registers and the sprite colour registers (`$D025`-`$D02E`)
now take effect while a sprite is being shifted out, following the
chip's sprite data sequencer as VICE's cycle-based VIC-II establishes
it. The data register shifts one bit per pixel under an expansion
flip-flop, so setting X-expand repeats the pixel being shown and
clearing it fetches the next at once. A multicolour pair is taken every
other pixel under a second flip-flop that a change of the multicolour
bit clears, so the pairs realign on the shape's odd bits (the sprite
split effect). The priority bit is read at every pixel, and a sprite
colour written mid-line changes from that pixel like the background
colours do.
- The three mode registers go through the sprite register journal in
`Vic2` with their own visibility points: priority and X-expand two
pixels before an X write is seen, multicolour one pixel before. The line
walk stores per-run flags (X-expand, multicolour, priority as the run
starts) and, when a journalled write changes one of the sprite's bits
inside the run, decodes the run pixel by pixel with a static decoder in
`Vic2SpriteManager`. Sprite-to-sprite collision masks use the decoded
pixels too.
- The sequencer generator shadows the ten sprite colour registers
through its write journal like the background colours, takes each run's
colours as the run starts, records colour changes that land inside a
run, and draws decoded or colour-split runs pixel by pixel. Plain runs
keep the existing fast path. The legacy generator's band model is
unchanged.
- Krestage 3 (Crest, 2007) is added to the Avalonia and browser Download
& Run lists and the docs table. It checks the chip for these effects
before it starts.
- Docs updated; seven new tests cover the decoder cases, the journal
timing and a colour split.

## Verification

- VICE spritesplit suite: 17 of 17 match the reference (16 were DIFF).
- Whole VICE harness, 274 tests on both models: no other status or
pixel-count change.
- Games A/B against the integration branch: identical for Commando (two
snapshots, D64 PAL and NTSC) and Giana Sisters (D64 PAL and NTSC).
- Krestage 3 passes its "NO VIC INSIDE" check and runs through 45000
frames; from its disk through the Download & Run path the intro figure
matches csdb's screenshot and the main screen follows.
- Demo A/B between the integration build and this one: For Your Sprites
Only now shows its per-line colour writes from the pixel they land on
inside a sprite row (the ss-mc-color rule), eight other sprite demos
render identically.
- Timing on For Your Sprites Only (300 + 600 frames, two rounds): level
within noise on every configuration.
- Whole solution builds; full test suite passes (the known wall-clock
timing test passes when rerun alone).
…d collisions from the line's pixels, fix the collision interrupt bits (#331)

## Summary

- **Sprite priority by the topmost sprite.** Where sprites overlap, the
chip shows the lowest-numbered sprite with an opaque pixel and lets that
sprite's priority bit alone decide against the foreground graphics
(VICE's spritepriorities test). Our two-layer drawing let a sprite in
front of the graphics show through a higher-priority sprite that was
behind them. The sequencer generator now composites each line's sprites
before writing them to the layers, on the lines where two sprites
overlap; lines without overlap keep the direct writes.
- **Sprite-to-background collisions from the line's pixels.** Derived at
the line's end from the sprite runs' opaque pixels against the
foreground pixels the graphics sequencer output (a set bit, or a 10/11
pair in multicolour; nothing while the vertical border flip-flop is
set), latched through
`IVic2SpriteManager.AddSpriteToBackgroundCollisions`. The old comparison
of the sprite shape with the character data under its start-of-line
position treated every set bit as foreground and every cell of a
multicolour screen as multicolour; it stays for the legacy generator and
the per-line-off paths.
- **Collision interrupt bits.** The flags now latch whether or not the
source is enabled in `$D01A` (the mask gates the CPU's interrupt line,
not the flag), and the two collision sources take their bits as the chip
has them: bit 1 sprite-background (IMBC), bit 2 sprite-sprite (IMMC).
They were swapped, so both collision interrupts used the wrong bits in
`$D019`/`$D01A`. Reading `$D01E`/`$D01F` clears the collisions and
leaves the flag.
- Docs updated; four new tests cover the topmost-sprite rule, the
collision definition in multicolour (01 pair none, 10 pair collides, a
single-colour cell in multicolour mode collides), and the flag latched
with the source disabled and kept across a register read.

## Verification

- VICE harness, 274 tests on both models: spritepriorities/test1 (1512
pixels before), sprite-gfx-hi-mc, sprite-gfx-mc-mc and
spritecollclear/bug2135 go from DIFF to MATCH. Nothing else changes
except vsp-tester and vsp-tester-ntsc: that test reads the idle byte's
address through the sprite-background collision register, never got a
value on the old approximation and ran into the harness timeout (counted
as a match for a test without a reference), and now reaches its verdict
and fails on the DMA-delay idle byte, which is not modelled yet.
- Games A/B against the integration branch: identical for Commando (two
snapshots, D64 PAL and NTSC) and Giana Sisters (D64 PAL and NTSC).
- Timing on For Your Sprites Only (300 + 600 frames, two rounds): the
per-line sprite path 2-5% slower (a vectorised foreground search before
any collision decoding and direct writes on lines without overlap keep
it there), every other configuration level.
- Whole solution builds; full test suite passes (the two wall-clock
timing tests pass when rerun alone).
…vents per model, take the VIC bank from CIA 2's port pins (#332)

## Summary

- **The display decision needs the enable bit.** The chip's cycle 58
check asks for the sprite's enable bit as it stands in that cycle, which
the VIC-II article's rule 4 omits. A sprite switched on for the two DMA
compares and off again before cycle 58 is fetched but not shown (VICE's
spriteenable 1, 2 and 4 test programs, which write the enable register
with a read-modify-write whose two writes land on the two compare
cycles).
- **The sprite event cycles per model.** On the 6567R8's 65-cycle line
the compares are in cycles 56 and 57 and the decision in 59, one later
than PAL. The offsets were PAL constants; they now follow sprite 0's
pointer cycle per model, as the bus stall model already did.
- **A late DMA start leaves sprite 0's first byte to the CPU.** Started
by the second compare, its fetch is two cycles on, one short of what BA
needs, so that byte reads $FF.
- **The display is cleared at the cycle-58 decision when the DMA is
off**, not in cycle 16 when the DMA ends, so a sprite whose Y is
rewritten to its last line restarts there and shows its first row again
(VICE's spriterestart premise).
- **A sprite 3-7 shown on the line its DMA starts** carries what its
fetch slot read while the DMA was off: $FF, the idle byte, $FF (VICE's
sb_sprite_fetch readme).
- **The VIC bank follows CIA 2 port A's pins.** A bit the direction
register makes an input floats high through its pull-up, so programs
that select the bank through `$DD02` got bank 3 from us instead of bank
0. The snapshot restore derives the bank the same way.
- **SpriteEnable sample** in the Avalonia and browser menus: three timed
cases on the SpriteX sample's line clock, showing the $FF byte on a
sprite enabled between the compares, an enabled-then-cleared sprite that
never shows, and the restart on the last line, on PAL and NTSC.
- Docs updated; seven new tests cover the rules, two DMA tests moved
their display assertion from cycle 16 to 58, and the snapshot round-trip
test sets the direction register as the KERNAL does.

## Verification

- VICE harness, 274 tests on both models: banking and spriteenable 1 and
2 on PAL and NTSC go from DIFF to MATCH, spriteenable 4 and the three
spritebug tests improve, nothing gets worse. 194 tests match.
- Games A/B against the integration branch: identical for Commando (two
snapshots, D64 PAL and NTSC) and Giana Sisters (D64 PAL and NTSC).
- Demo A/B: eleven demos identical; Chars Sucks, which selects its bank
through `$DD02`, now shows its sprite logo and scroller (its ghost-byte
shading remains a follow-up).
- The sample's write cycles were verified with a register trace on both
models (line 212 index 54/55, line 208 index 55-56/56-57, line 233 index
55-56/56-57) and the run dumps (sprite 0's first byte $FF, sprite 1
clean, sprite 2 never displayed, sprite 1 restarted with row 0).
- Whole solution builds; full test suite passes (the two wall-clock
timing tests pass when rerun alone).

- Timing (For Your Sprites Only, 600 frames, three repeats, two
alternating rounds, MacBook Air M1): no measurable change against the
integration branch in any renderer configuration. Sequencer with
per-line sprites 1297-1322 vs 1310-1332 us/frame, the other four
configurations likewise within the round-to-round noise.
…8FF, fix the VICE harness's NTSC names and frame compared (#333)
…es (#338)

## What

The Y-expansion flip-flop of a sprite is inverted while its `$D017` bit
is set. The VIC-II article puts that inversion in the first phase of
cycle 55; on the chip it is a cycle later: a bit set by a write in cycle
55 is still inverted (the row is held, the sprite stretcher case), a bit
set in cycle 56 is not. We held the row only for writes up to cycle 54.

Established by VICE's `spritecrunch2` test programs, which clear and set
`$D017` on every line and move the second write a cycle later every
eight lines.

## How

`Vic2`: the inversion moves from the sprite event of the first DMA
compare (cycle 55) to that of the second (cycle 56), before the compare.
Sprites the first compare has just started are left out: their flip-flop
was cleared there, and a Y-expanded sprite still shows its first row
twice.

## Verification

- VICE test programs (274 runs, PAL and NTSC): 245 → 250 match. Newly
matching: `spritecrunch/spritecrunch2-25` to `-29`. No other result
changed, including the other nineteen sprite suites.
- New unit test
`The_flip_flop_inversion_sees_a_y_expand_bit_set_in_cycle_55`: set in
cycle 55 holds the row, in cycle 56 it does not.
- Games (Commando, Giana Sisters; PAL, NTSC, snapshots) and fourteen
demos at three frame counts render byte-identically before and after.
- Performance (MacBook Air M1, For Your Sprites Only): unchanged in
every configuration (sequencer with per-line sprites 1347/1320 →
1342/1318 µs/frame).
…d multicolour pair as the chip does (#339)

## What

Two details of a sprite drawn on the line its own data fetch happens,
each established by one of VICE's test programs.

**The bytes a fetch slot reads with the DMA off.** A sprite 3-7 with X
at or beyond `$164` is displayed on the line its compare starts the DMA,
with what its fetch slot at the line's start read while the DMA was
still off. That was `$FF`, the idle byte, `$FF`, with the idle byte read
when the line ended. Now:

- the two bytes of the CPU's half of the slot's cycles are the byte of a
VIC-II register access (read or write) the CPU makes in that cycle,
`$FF` when it makes none;
- the idle byte is the one in place as the line begins.

`sb_sprite_fetch/sbsprf24-164` places three bytes of its own this way
(`STA $D000,Y` timed so its dummy read and its write fall in sprite 6's
slot, and an idle byte it restores a few cycles later).

**A multicolour pair caught by the halt.** A sprite still shifting when
its own fetch begins repeats its last pixel for seven pixels. When a
multicolour pair is taken in the last pixel before the halt, that pixel
and its repeats show only the pair's high bit, as a standard pixel
would: `01` shows nothing, `11` the sprite colour.
`spritefetchbug/test-136-2a` shows it; the X positions its readme lists
are exactly that alignment, so the rule is kept that narrow.

## How

- `Vic2`: register accesses in a line's first ten cycles (the slots of
sprites 3-7) record their byte by cycle, reset per line only when
something was recorded; the idle byte is captured at the line's start
when an enabled sprite 3-7 with its DMA off has its Y on the line. The
run derivation builds the slot's row from those.
- `Vic2SpriteManager.DecodeSpriteRun`: a pair taken at the pixel before
the halt keeps only its high bit from that pixel on. The run derivation
sends such runs through the decoded path, so drawing and collisions both
see it.

## Verification

- VICE test programs (274 runs, PAL and NTSC): 250 → 252 match. Newly
matching: `sb_sprite_fetch/sbsprf24-164`, `spritefetchbug/test-136-2a`.
No other result changed.
- New unit tests: the slot's bytes from a register read and write with
the idle byte of that time; the pair rule (`11` → sprite colour, `01` →
nothing, a pair taken a pixel earlier unchanged).
- Games (Commando, Giana Sisters; PAL, NTSC, snapshots) and fourteen
demos at three frame counts render byte-identically before and after.
- Performance (MacBook Air M1, For Your Sprites Only): unchanged in
every configuration (sequencer with per-line sprites 1288/1305 →
1274/1281 µs/frame).
#340)

## What

A sprite enabled (or moved onto the raster line) after the line's two
DMA compares, in cycles 55 and 56, does not start fetching on that line.
The VIC-II model already had that right, but the bus stall model still
gave such a sprite a BA-low window, so a CPU read near the end of the
line was held for up to three cycles for a fetch that never happens.

## Why it happened

`Vic2BusStalls.BuildWindows` counts a sprite that is about to start
(enabled, DMA off, Y on the line) among the fetching ones, so that
sprite 0's window, which begins in the same cycle as the first compare,
is seen before that compare has run. The prediction was used for the
whole line. It is now used only until the second compare has been made;
from there the VIC-II's DMA state is final for the line.

## Verification

- VICE test programs (274 runs, PAL and NTSC): 252 → 254 match. Newly
matching: `spriteenable/spriteenable4` and `spriteenable4_ntsc`, which
enable sprites 0-2 in cycle 57 of the line their Y names and draw a
timing bar with the code that follows. No other result changed; the
sprite and DMA suites are 129 of 129.
- New unit test
`A_sprite_enabled_after_the_compares_takes_no_bus_on_that_line`: enabled
in cycle 55, a read in cycle 57 waits three cycles; enabled in cycle 57,
none.
- Games (Commando, Giana Sisters; PAL, NTSC, snapshots) and fourteen
demos at three frame counts render byte-identically before and after.
- Performance (MacBook Air M1, For Your Sprites Only): unchanged in
every configuration.
## What

In the first phase of every cycle the VIC-II reads memory, and that byte
stays on the data bus into the second phase unless something else drives
it. Two kinds of CPU read now see it:

- **The I/O 1 and I/O 2 areas (`$DE00`-`$DFFF`)** with no cartridge
answering the address. They returned whatever had last been written
there (I/O storage), which a real C64 does not have.
- **Colour RAM**, whose chip drives only the data bus's low four bits:
the high four now come from that byte instead of being 0.

Which byte depends on the cycle of the read:

| Cycles (PAL, 1-based) | First-phase access | Byte |
|---|---|---|
| 58, 60, 62, 1, 3, 5, 7, 9 | pointer access of sprite 0-7 | the pointer
|
| the cycle after each | sprite data slot | middle byte of the row if
the sprite's DMA is on, else `$3FFF` |
| 11-15 | refresh | `$3F00` + refresh counter (reset to `$FF` in line 0,
counted down per access) |
| 16-55 | graphics | idle state `$3FFF` (`$39FF` with ECM); display
state the character or bitmap data |
| the rest | idle | `$3FFF` |

On the 6567R8 (NTSC, 65 cycles) the sprite accesses start a cycle later
(sprite 0 in cycle 59).

## How

`Vic2.FirstPhaseBusByte()` brings the VIC-II to the cycle of the access
in progress and derives the byte from the cycle, the sprite DMA state
and the display state. The cartridge slot uses it as its fallback reader
for `$DE00`-`$DFFF` (writes still go to I/O storage), and `ColorRAMLoad`
takes its high four bits from it.

Five cartridge and SwiftLink unit tests expected an unanswered I/O read
(a write-only register, a detached SwiftLink) to give back the stored
value; they now expect the bus byte. As on hardware, a program probing
for an REU by writing a register and reading it back no longer finds one
that is not there.

## Verification

- VICE test programs (274 runs, PAL and NTSC): 254 → 256 match. Newly
matching: `phi1timing/phi1timing` and `phi1timing_ntsc`, which read
`$DEAD` on every cycle of a line and compare each byte with the expected
access. `colorram/test` (reads colour RAM, masks to the low four bits)
still passes. No other result changed.
- New `Vic2FirstPhaseBusByteTests` (17 cases): pointer, slot, refresh,
graphics and idle accesses by cycle, ECM for idle graphics accesses
only, the refresh counter, a fetching sprite's middle byte,
display-state character data, and colour RAM's high bits.
- Games (Commando, Giana Sisters; PAL, NTSC, snapshots) and fourteen
demos at three frame counts render byte-identically before and after.
- Performance (MacBook Air M1): For Your Sprites Only unchanged
(sequencer with per-line sprites 1357/1281 → 1349/1280 µs/frame). A
worst-case loop doing nothing but colour RAM reads (~3,400 per frame) is
228 → 271 µs/frame without a renderer and 637 → 682 with the sequencer:
about 13 ns per colour RAM read, half of it bringing the VIC-II to the
read's cycle as register reads already do. A scroller copying all of
colour RAM once per frame pays about 13 µs.
…/O visible (#342)

## What

Two changes, both found through VICE VIC-II test programs that did not
reach their real end.

### LXA and ANE

LXA (`$AB`, `LAX #imm`) and ANE (`$8B`, `XAA #imm`) were not implemented
in any compatibility profile. The CPU ran such a byte as a one-byte
instruction and then executed its operand as the next opcode, so a
program using them derailed. `flibug`'s FLI displayer (from Black Mail's
FLI Graph editor) runs `LXA #0`; the operand `$00` was executed as `BRK`
and the program ended in BASIC's warm start.

Both are "unstable" opcodes: their result ORs a chip-specific value into
A before the AND.

- `LXA #imm`: `A = X = (A | $EE) & imm`
- `ANE #imm`: `A = (A | $EE) & X & imm`

`$EE` is the common value, and it is settled by the SingleStepTests 6502
corpus already in the test fixtures: all 20 LXA and all 20 ANE vectors
match `$EE` (6 and 15 of them match `$FF`). The two opcodes are removed
from the corpus test's known deviations, so their results, flags and
cycles are now asserted.

They are in the `StableUnofficial` profile, the C64's and VIC-20's
default: running them with the common value is closer to any real chip
than running them as the wrong instruction. The profile's help text in
the Avalonia config dialog and the CPU library docs say so.
`OpCodeId.LXA_I` and `OpCodeId.ANE_I` are added.

The other unstable opcodes (SHA, SHX, SHY, TAS) stay unimplemented:
their result also depends on whether the video chip takes the bus in a
particular cycle, and they are a separate piece of work.

### The VICE harness's exit register

The harness took any write to `$D7FF` as the program's exit code, in all
32 memory configurations. The testbench register is in the I/O area, so
a program that clears memory with I/O banked out (`colorfetchbug/main`
does) was stopped by a write to RAM. The harness now hooks `$D7FF` only
in configurations where I/O is visible, from a new `C64.IsIOVisible`,
and runs the programs with the `FullUnofficial` profile, since they are
written for the real chip.

## Verification

- VICE VIC-II test programs (274 runs, PAL and NTSC): 256 → 258 match.
`flibug/blackmail-ee` and `blackmail-fixed` now match their references
pixel for pixel (the `-ee` variant depends on the constant).
`colorfetchbug/main` now runs to its exit and differs in 7 pixels where
its reference, a VICE screenshot, lacks the idle byte hardware shows
when a bad line starts mid-line. `vspbug/vsp_bug` now really exits `$00`
(it was stopped with `$5A` before). No other result changed.
- VICE CIA test programs (104 runs): identical before and after.
- CPU tests: the corpus now asserts LXA and ANE; profile tests cover
where they are defined.
- With the C64's default settings the `flibug` displayer now runs
instead of dropping to `READY.`.
- Games (Commando, Giana Sisters; PAL, NTSC, snapshots) and fourteen
demos at three frame counts render byte-identically before and after,
with the new default profile.
- Performance (MacBook Air M1, For Your Sprites Only): unchanged.
…add Robot - Not Human (#344)

## Summary

- The CIA timer never shows 0 while counting: the count from 1 is the
underflow, the latch shows for two cycles, a latch of 0 counts like 1, a
latch written in the underflow cycle is the one reloaded, a start from a
stopped counter of 0 underflows two cycles later, a force load leaves
the interrupt flag alone, and a one-shot timer stops in its underflow
cycle.
- The interrupt output follows the underflow a cycle later and the CPU
sees it a cycle after that, so an interrupt control read in between
still keeps an IRQ from being taken but not an NMI.
- Enabling a source whose flag is already set drives the interrupt
output the same way.
- Snapshot restore and freeze drop the pending output event with the
other pipeline events.
- Tests that encoded the old visible-zero and latch-zero behaviour are
rewritten; a test covers the mask-enable rule. The CIA section of
`docs/systems/c64/libraries.md` describes the rules.
- Robot - Not Human (Cycleburner and Pal, 2026) is added to the Avalonia
and browser Download & Run lists, with a row in
`compatible-programs.md`. It starts a one-shot timer from 0 and enables
its NMI afterwards, so it runs only with the mask-enable rule.

## Verification

- VICE CIA test programs: 62 of 104 match, up from 53. Newly matching:
`reload0a/b`, `irqdelay`, `irqdelay-oneshot`, `irqdelay-cia1-4-old`,
`timerbasics` `timer`/`timer_new`/`timer_test1`, `ciavarious`
`cia1`/`cia2`/`cia5`. `irqdelay-new` and `irqdelay-oneshot-new` (6526A
timing) matched before only because the output came a cycle early and no
longer do.
- VICE VIC-II test programs unchanged at 258 of 274.
- Commando, Giana Sisters (disk and snapshots), 14 demos and the
CiaSyncedSplit sample render identically to the integration branch;
Robot - Not Human is the only program whose output changes. The Expert
cartridge's freeze enters its monitor as before.
- Unit tests pass (the two wall-clock `ElapsedMillisecondsTimedStat`
tests fail only under load and pass alone).
- Frame benchmark within noise; the CIA timer micro-benchmark's
frequent-underflow case is ~8% slower (23.5 → 25.5 ns) from the separate
interrupt-output event.
…ed after the CPU's poll (#345)

## Summary

- The CPU records the cycle the IRQ line was released on
(`CPUInterrupts.IRQReleasedAtBusCycle`, `IRQWasActiveAt`) and still
takes an interrupt it sampled active before that cycle. A CIA interrupt
control read that releases the output in an instruction's last cycle is
too late to cancel the interrupt; the handler then finds the register
already cleared. The VIC-II's raster acknowledge still passes no cycle,
as before.
- Timer B's flag shows in the interrupt control register in the
underflow cycle, like timer A's, instead of a cycle earlier. What the
earlier reading of the reference data saw was a quirk of the 6526: a
read in the cycle before timer B's underflow loses that flag, which is
never set, while the interrupt output still follows a cycle after the
underflow, so a handler reads only the interrupt bit. Timer A's flag
survives such a read.
- A start leaves the interrupt flag alone. The clear on start dated from
the first CIA implementation and had no reference behind it; the
force-load half had already gone in #344.
- A read in the cycle after a read that took an enabled source's flag
shows the interrupt bit set, without the output being driven (`inc
$dd0d,x` reads the register in two consecutive cycles).
- `CiaIRQ` keeps its enable and flag state as two bitmasks at the
register's bit positions, the same representation as `CPUInterrupts`, so
a register read returns the mask itself; the CPU's `InterruptSource`
handles are cached so raising and releasing the lines is no longer a
name lookup. Public surface unchanged.
- Tests cover the released-line rule on the CPU and the three CIA rules;
the CIA paragraph of `docs/systems/c64/libraries.md` describes them.

## Verification

- VICE CIA test programs: 68 of 104 match, up from 62. Newly matching:
`cia-timer-oldcias`, `cia-timer-alt-oldcias` (the screen equals
`dump-oldcia.bin`), `ciavarious` `cia3`/`cia3a`/`cia4`, `dd0dtest` (all
19 sub-tests). No losses; the remaining old-chip failures are the
cascade/CNT, shift-register, PB6/PB7 and TOD `hammerfist` programs.
- VICE VIC-II test programs unchanged at 258 of 274, the same 16.
- 30 one-file demos render hash-identically over 1500 frames against the
integration branch.
- Unit tests pass (the two wall-clock `ElapsedMillisecondsTimedStat`
tests fail only under load and pass alone).
- Frame probe at parity (0.62 vs 0.61 ms/frame on a timer-heavy program
with load < 3); the interrupt poll costs one extra comparison only when
the line is inactive and was released.
…terrupt state as bitmasks (#346)

## Summary

- `Vic2IRQ.ClearTrigger` (a `$D019` write) and `Disable` (a `$D01A`
write) release the IRQ line on the write's cycle. An acknowledge in an
instruction's last cycle is too late to stop an interrupt the CPU
sampled at its second-to-last cycle; the handler is entered with the
register already cleared. This is the rule #345 gave the CIA's interrupt
control read, applied to the VIC-II. A test in
`C64DeviceAccessTimingTests` covers it.
- `Vic2IRQ` keeps its enable and latch state as two bitmasks at the
registers' bit positions, the same representation as `CPUInterrupts` and
`CiaIRQ`, with the CPU's `InterruptSource` handles cached per CPU
instance. The `$D019`/`$D01A` load and store handlers read and write the
masks directly instead of walking the enum with dictionary lookups on
every access. Public surface unchanged.
- Comments only: the `ColorMaps` tables are described as the Godot and
C64HQ palettes distributed with VICE (values untouched); the D64
free-block count's comment says it is 664 minus the directory's blocks
and does not read the BAM.

## Verification

- VICE VIC-II test programs unchanged at 258 of 274, the same 16.
- 30 one-file demos render hash-identically over 1500 frames against the
integration branch.
- Systems tests pass (Commodore64 and snapshot suites, 639).
- Frame probe on Party Elk 2 (a `$D019` write every raster line): 1.14 →
1.04 ms/frame, from the register handlers no longer allocating
(`Enum.GetValues`) or looking up dictionaries per access.
…e on the cycle after the port write (#347)

## Summary

- `ARR` ($6B) moves from `ExperimentalUnofficial` to `StableUnofficial`,
the C64's default profile. Its binary-mode result is deterministic on
every NMOS 6502 and the single-step vector corpus asserts it; `LAS`,
whose result depends on the bus, stays experimental. Party Elk 2 (Booze
Design, 2023) computes its FPP tables with `ARR #$FE` in 58 unrolled
places and derailed under the default profile after its first part,
showing the "scattered red and blue fragments" of the open bug; Rainbow
Connection (Hein, 2023) uses `ARR #$80` for its star animation and
showed three frozen stars. Profile help text and
`docs/libraries/core/dotnet6502.md` updated, `OpCodeInfoTests` adjusted.
- A VIC-II bank change from CIA 2's port (`$DD00` or `$DD02`) reaches
the chip's fetches in the cycle after the write, as a register write
does: the chip fetches in a cycle's first phase and the CPU writes in
its second, so the fetch of the write's own cycle and those of the
instruction's earlier cycles still read the old bank. Party Elk 2 writes
`$DD02` in cycle 53 of every scroller line, and on the lines where that
moved the bank (every ~29 lines) we fetched columns 33-36 from the new
bank: broken rows at one fixed screen column as letters scrolled
through. `Vic2.SetVIC2BankFromPortWrite` dates the change and
`CatchUpTo` lands it part way through a catch-up, with two writes
allowed in flight (a read-modify-write); the direct `SetVIC2Bank`
(tests, snapshot restore) stays immediate. A test in
`C64VideoMemoryWriteTimingTests` uses the idle byte in two banks and a
port or direction-register write timed to cycle 30.
`docs/systems/c64/libraries.md` states the rule.
- Party Elk 2 is added to the Avalonia and browser Download & Run lists,
with a row in `compatible-programs.md`.

## Verification

- VICE VIC-II test programs unchanged at 258 of 274, CIA at 68 of 104.
- Of the 30 cached one-file demos, run under the default profile against
the integration branch, 28 render hash-identically over 1500 frames;
Party Elk 2 and Rainbow Connection are the two the change repairs.
- Unit tests pass (the wall-clock `ElapsedMillisecondsTimedStat` tests
fail only under load).
- `C64ExecuteFrameBenchmark` (`--job short`, alternating): a first shape
with the pending-bank loop inline in `CatchUpTo` (per instruction and
per VIC register access) cost 2-4% in every scenario; the loop is now in
a `NoInlining` method behind one int test, and the benchmark is at
parity (CoreOnly 277.8 → 277.9 µs, MixedVisibleSprites 283.4 → 281.9,
RenderOnly 428.5 → 432.0 / 455.6 → 441.4, RenderAndAudio 724 → 692 / 717
→ 711).
…l commands (#348)

## Summary

- **Debug values.** A system can expose named values to debuggers
through `IDebugValueSource` (`Highbyte.DotNet6502.Systems.Debugger`).
The C64 exposes its VIC-II position: `RASTER`, `CYCLE` (the chip's
1-based numbering, 1–63 PAL / 1–65 NTSC), `FRAMECYCLE` and `FRAME` (a
new `Vic2.FrameCount`). The chip is brought up to `CPU.BusCycles` on
read, so a CPU-only monitor single step shows the right position.
`IDebugValueSource.FormatDebugValuesLine()` is the one formatter both
the monitor and the hosts' status area use, so they cannot drift.
- **Conditions.** `BreakpointConditionEvaluator` accepts those names as
operands (registers and flags still win), computes in `long`, and gains
`TryParse` so a mistyped condition is rejected up front instead of
stopping on every hit. `&&` binds tighter than `||` (was already so; now
tested and documented).
- **Run until.** `DebuggerBreakpointEvaluator.RunUntilCondition` is
checked before every instruction regardless of PC and stops once
(`ExecEvaluatorTriggerReasonType.RunUntilCondition`); a pause or a stop
for another reason drops it.
- **Monitor.** `r` prints a `VIC-II: RASTER=… CYCLE=… FRAMECYCLE=…
FRAME=…` line; `gu <condition>` continues until a condition holds; the
C64 adds `gr <line> [cycle]`. The C64's first `SystemInfo` row (shown by
the Avalonia, SadConsole, SilkNet, WASM and terminal hosts) is now that
same line, with the banks moved to the model row.
- **VS Code.** A `VIC-II` scope in Variables (read-only), evaluation of
the names on hover, in Watch and in the Debug Console, conditions such
as `RASTER == 100 && CYCLE >= 20`, and the Debug Console commands `run
until <condition>` and `run raster <line> [cycle]`. Run targets come
from `IDebugValueSource.TryBuildRunUntilCondition`, so the adapter stays
system-agnostic. `C64.BuildRunUntilRasterCondition` builds a
`FRAME`/`FRAMECYCLE` condition so a target in a line's last cycles is
not skipped by an instruction straddling the line end (a plain `RASTER
==`/`CYCLE >=` check would wait a frame). Stops are
instruction-granular: the position shown is where the next instruction
starts, up to one instruction past the target. Extension `CHANGELOG.md`
`[Unreleased]`, `README.md` and `DEBUGGING.md` (new "System values" and
"run" sections) updated; a "Run to raster line…" command-palette entry
is deferred.
- **Two pre-existing debug adapter races fixed**, found because `run
until` makes a stop happen microseconds after a resume: `BaseTransport`
had no write lock, so a `stopped` event sent from the run loop could
interleave its bytes with a response on the wire; and the shared
debug-log `StreamWriter` was written from two threads, throwing inside
`LogSafe` and silently dropping the stopped event. Both now serialize on
the writer.

## Verification

- 0 warnings; all test projects green (3,150 tests). New:
`BreakpointConditionEvaluatorTests` (syntax, precedence, `long`,
`TryParse`), `DebuggerBreakpointEvaluatorRunUntilTests`,
`C64DebugValuesTests` (values on PAL and NTSC, frame count, CPU-only
step, raster conditions ahead/passed/current/straddling, argument
validation, `SystemInfo` row), `C64MonitorGoRasterCommandTests`,
`RunUntilAndSystemValuesCommandTests` (monitor `r`/`gu`), and
`DebugAdapterSystemValuesTests`, which drives `DebugAdapterLogic` over
DAP with a C64 behind it (scope, variables, hover, `run raster`, `run
until`). `Highbyte.DotNet6502.Systems.Tests` now references the
DebugAdapter project for that.
- Manual: DAP smoke over stdio against the console adapter, and an
end-to-end session from the VS Code extension against the Avalonia host
with `smooth_scroller_and_raster2.asm` (scope, Watch, conditions, `run
raster`, `run until`), plus the Avalonia monitor's `r`/`gr`/`gu` and
status lines.
- `C64ExecuteFrameBenchmark` (`--job short`, base/branch/base/branch,
load 1.5–2.9): parity within noise — CoreOnly 273/278 vs 273/279 µs,
RenderOnly 423–442 vs 426–441, AudioOnly 420–446 vs 428–453,
RenderAndAudio 679–708 vs 686–713. The benchmark runs no evaluator, so
the only hot-path change it exercises is `FrameCount++` once per frame;
with a debugger evaluator installed the added cost per instruction is
one null check.
…bench harness (#349)

Testbench tooling for `tools/vice-testprogs`, the console harness that
runs VICE's test programs against the C64 emulation.

## What changes

- **SYS expressions.** The BASIC stub's `SYS` argument may be an
expression, as in the 64doc tests (`SYS PEEK(43)+256*PEEK(44)+26`); the
harness parsed only `SYS n` / `SYS(n)` and those tests never started. A
small evaluator now handles numbers, `PEEK(...)`, parentheses and `+ -
*` on the machine's own memory. All seven `CPU/64doc` tests now run (and
pass).
- **`--exclude <substring>`**, the complement of `--filter`.
- **`--testlist <c64-testlist.in>`** reads VICE's own list of C64 tests,
which makes the verdicts exact:
- tests marked `expect:timeout` (the CPU jam tests) or `expect:error`
pass on that outcome and fail otherwise;
  - `interactive` / `analyzer` entries are skipped and counted;
- each test runs for its listed cycle budget instead of `--frames`
(`CPU/decimalmode/scanner` needs ~36 000 frames; the default budget was
600).
- **Three-way verdict** `pass` / `fail` / `timeout` in `results.md` and
the console summary. Before, a program that never wrote its exit code
counted as a match, so a hang and an expected hang were
indistinguishable from a real pass.
- `docs/home/development.md`: the harness section updated for the above.

## Verified locally

- `CPU/64doc/*` (7): timeout → all `$00`. `CPU/decimalmode/scanner`:
`$00` after 36 121 frames.
- With the test list: `cpujam02` → `pass (expected timeout)` at its
280-frame budget; `ciaports` skipped as interactive; `irqdma/test1` runs
its 22 893-frame budget overriding `--frames`; the ambiguous
`scanner.prg` (also under `drive/`) resolves to `CPU/decimalmode`.
- No change in outcome for the other `SYS` forms in the suites (`SYS
2064`, `SYS(2080):…`, `SYS2069 /…/`), spot-checked against the previous
run's results.
…age-crossing and RDY behaviour (#350)

The five NMOS "unstable" indexed stores — `SHA` ($93 `(zp),Y`, $9F
`abs,Y`), `SHX` ($9E `abs,Y`), `SHY` ($9C `abs,X`) and `TAS`/`SHS` ($9B
`abs,Y`) — were not implemented in any compatibility profile. The CPU
ran each byte as a one-byte instruction and executed its operand bytes
as code, which derails any program using them: every plain program of
the Wolfgang Lorenz suite's `trap*` family died at its first `SHA`, and
VICE's `CPU/sha`, `shs` and `shxy` programs failed or hung.

## Behaviour

One composition, `ComposeUnstableStore`, binds all five from
`StableUnofficial` upwards (the C64 and VIC-20 default, as decided
earlier for `LXA`/`ANE`: a wrong-length NOP is worse than any
approximation):

- The value stored is the register (`A & X`, `X`, `Y`, or `A & X` for
`TAS`, which sets `SP` to it first) ANDed with the high byte of the
un-indexed address plus one.
- When the index carries into the high byte, that ANDed value becomes
the high byte of the address written to.
- When RDY is low in the cycle before the write — the dummy read is
stalled, on the C64 by the VIC-II taking the bus — the AND drops out and
the register value alone is stored; the address is corrupted the same
way either way. The composition reads a new `internal
CPU.StallCyclesInProgress` around the dummy fetch; nothing on the hot
path of other instructions changes.

Sources: the "NMOS 6510 Unintended Opcodes" document and the readmes of
VICE's `testprogs/CPU/sha` and `shxy`; the SingleStepTests corpus
settles the RDY-high cases including page crossings. No VICE source code
consulted.

## Verification

- **SingleStepTests corpus:** the five bytes removed from the known
deviations; all 100 vectors (52 with a page crossing) are now asserted
and pass.
- **New `UnstableStore_test`:** the RDY-low drop-out in the right cycle
(and not one cycle earlier), address corruption with the AND dropped,
`TAS`'s stack pointer, profile gating, cycle counts.
- **VICE test programs** (`--testlist`): `CPU/sha`, `shs`, `shxy` (all
31, including the DMA-timing programs `*3`/`*4`/`*5` that previously
hung), `asap/cpu_shx` and `Acid800/cpu_illegal` all pass. Lorenz suite
279/295 (was 245): every `trap*`, `shaay`, `shaiy`, `shsay`, `shxay`,
`shyax` and `irq` pass. CPU + interrupts 125/142 (was 108). The
remaining failures are unrelated (CIA timing, NMI/IRQ interplay, 6510
port, and `CPU/ane`, which checks the ANE magic constant under RDY and,
per its readme, also fails on some real C64s).
- Full test suite green; two tests that used "$9C undefined on NMOS" as
a model probe now probe the mnemonic.
- `C64ExecuteFrameBenchmark` at parity against the base branch.

## Also

- Testbench harness: a program listed under several directories sharing
a path component (`CPU/asap` vs `CPU/Acid800` for suite `CPU_asap`) is
now matched to the most specific one, so it gets its cycle budget.
- Docs: profile table, corpus note and the Avalonia profile description
updated.
…351)

A 6510 port bit that nothing drives keeps, while configured as an input,
the value it was last driven to as an output: the pin behaves like a
small capacitor that the output charged and that the input then reads
back (for a few hundred milliseconds on the real chip). That applies to
bits 6–7, which have no pins, and on the C64 to bit 3, the cassette
write line, when no datasette is attached. The port model read the live
latch for bits 6–7 and a hard 0 for bit 3, which two test programs
catch:

- VICE `CPU/cpuport/test1` (bit 7): after switching the bit to input, a
write to `$01` must not change what reads back.
- Wolfgang Lorenz `cpuport` (bit 3): `$FF` to both registers, then all
inputs, then `$FF` to `$01` must read `$DF`; ours read `$D7`.

## Change

`Cpu6510Port` samples the latch into a held charge for every
output-configured bit whenever the direction register is written, so a
bit turning into an input holds what it last drove. Reading an
input-configured bit that is unimplemented (6–7) or listed in the new
board-set `FloatingLinesMask` returns that charge; driven inputs still
read `ExternalInputLevels`. The C64 sets the mask to bit 3 (no
datasette; a datasette device would drive the line through
`ExternalInputLevels` instead). No decay is modelled.
`SetState`/snapshot restore seed the charge from the data register, so
the snapshot format is unchanged; `Clone` copies it.

## Verification

- Unit tests: the `test1` sequence on bit 7 step by step, the Lorenz
sequence with C64 wiring, and clone; the existing port tests pass
unchanged.
- `CPU/cpuport/test1` and Lorenz `cpuport` pass. Full Lorenz + banking
suite: only `cpuport` changed (280/295).
- Full VICII (274) and CIA (104) test-program runs: exit codes and pixel
counts identical to the base branch, compared per test.
- Test suites green; `C64ExecuteFrameBenchmark` at parity (the port read
is not on any per-instruction path).
)

The CPU samples the IRQ line during an instruction's second-to-last
cycle. A device release dated to that very cycle was treated as
"released before the sample", so the interrupt was not taken. On
hardware the CPU samples the line before the register access that
releases it lands at the end of the cycle, so the interrupt is still
taken. The only case that shows it is an instruction whose
second-to-last cycle is itself the releasing access: a read-modify-write
of an interrupt register whose dummy write (the old value written back)
acknowledges, as `INC $D019` / `ASL $D019` do when the raster flag is
set.

Worked out from VICE's `interrupts/irq-ackn-bug/irq-ack-vicii`, which
times `sta`, `inc`, `asl` and `lda $D019` at six one-cycle offsets
around a raster interrupt: the `inc` and `asl` columns at the two latest
offsets showed no interrupt in ours where the reference shows one. The
`$D019` write path itself was already right (the dummy write does
acknowledge; the reference rows are consistent with that).

## Change

`CPUInterrupts.IRQWasActiveAt`: a release in the sampled cycle no longer
cancels the interrupt (`>=` instead of `>`; a release with no cycle
given still counts as before any poll). One comparison on the interrupt
poll; `C64ExecuteFrameBenchmark` at parity.

## Verification

- `CPUInterruptSamplingCycleTests`: the release-after-poll theory now
covers the poll cycle itself; a new `C64DeviceAccessTimingTests` case
runs `ASL $D019` with the raster flag set and checks the interrupt is
taken although the dummy write acknowledged.
- `irq-ack-vicii`: the raster row now matches the reference character
for character. The program still exits `$FF` on its second half
(sprite-to-sprite collision interrupt timing: our collision flag is
raised later than hardware, as collisions are latched per line rather
than at the pixel) — a separate VIC-II item, not touched here.
- `interrupts/cia-int/cia-int-irq` passes as a consequence (same rule,
CIA register).
- Full VICII (274), CIA (104), Lorenz + banking (295) and CPU +
interrupts (142) test-program runs: identical to the base branch per
test apart from the two improvements. Test suites green.

## Also

Testbench harness: each test now also gets a `<test>.screen.txt` with
the text screen as ASCII, for reading a program's report without a
reference picture.
…errupt output until its register is read (#353)

Three findings, each pinned by a hardware dump or reference that ships
with the VICE test programs, together making
`interrupts/irqnmi/irqnmi-old`,
`interrupts/branchquirk/branchquirk-nmiold`, Wolfgang Lorenz's `nmi` and
Acid800's `cpu_bugs` pass.

## 1. An NMI hijacks an IRQ or BRK entry sequence (CPU)

An NMI that arrives by the 4th cycle of a 7-cycle IRQ or BRK sequence
takes the sequence over: it completes with the NMI vector and the stack
frame already pushed (a BRK's keeps B set, which Lorenz's `nmi` and
`cpu_bugs` check explicitly). A sequence does not poll the interrupt
lines at its end, so an NMI arriving later is taken after the handler's
first instruction. Neither was modelled: the NMI was taken as a second
sequence right after the first, which `irqnmi-old`'s dump shows as the
NMI handler entering 7 cycles late over a 7-offset span.

Implemented as a "hijackable entry" state: an IRQ sequence or BRK leaves
its poll point at the vector-decision cycle and, at the next boundary
(when the devices have been caught up), a pending NMI dated by that
cycle switches PC to the NMI vector instead of starting a sequence. BRK
carries a new `OpCodeDescriptor.IsInterruptEntry` marker. The 4th-cycle
threshold is calibrated against the dump (offsets 0–6 hijack, 7 and 8 do
not).

## 2. The CIA holds its interrupt output until the interrupt control
register is read (CIA)

The CIA's IRQ sources were registered as auto-acknowledging, so the CPU
dropped the line when it serviced the interrupt. The 6526 holds the line
until the ICR is read. With the hijack in place this mattered: the
hijacked IRQ was lost instead of being taken when the NMI handler
returned (the IRQ columns of `irqnmi-old` went blank). Every test
program and the KERNAL read the ICR, so nothing else moved.

## 3. Timer latches reset to `$FFFF` (CIA)

`cpu_bugs` writes a stopped timer's high byte (which loads the counter
from the latch) before its low byte. With the latch reset to 0 the
counter became 0, and the force-load-and-start then underflowed at once,
3 cycles after the start instead of 7. The timer latches are the one
register the 6526 leaves at all ones on reset. This also fixed the
single differing cell in `branchquirk-nmiold` (its very first timer
start).

## No regressions

Every test-program suite re-run on this branch and compared program by
program with the previous branch head: VICII 274 (PAL + NTSC) and CIA
104 identical; Lorenz + banking 295 identical apart from `nmi` (fail →
pass); CPU + interrupts 142 identical apart from `irqnmi-old`,
`branchquirk-nmiold` and `cpu_bugs` (fail → pass). Unit-test suites
green. `C64ExecuteFrameBenchmark` three base/branch pairs within
run-to-run noise (the one per-instruction addition is a store and a
branch in the poll-point bookkeeping); the RenderOnly/None scenario read
~3% higher in all three pairs without a path that explains it, noted
here rather than hidden.

New tests: the hijack window on an IRQ sequence and on BRK, the frame
contents in both cases, the IRQ being taken after the hijacking NMI's
handler returns, and the absence of an end-of-sequence poll.

## Also

- `interrupts/cia-int/cia-int-nmi` still fails on its last column: that
needs the ICR mask-disable write to take effect a cycle late, symmetric
to the enable delay already modelled — CIA interrupt-control timing, not
touched here.
- Testbench harness: each test also writes `<test>.screen.bin`, the raw
text screen, for byte-by-byte comparison with the reference dumps.
- Docs: the interrupt-sampling paragraph of the core library page
updated.
…able miss timer B's output due in the write cycle (#354)

Two corrections to how a write to the CIA's interrupt control register
reaches the chip's interrupt output, from Wolfgang Lorenz's `imr` and
VICE's `interrupts/cia-int` test programs.

## Enable: seen by the CPU two cycles after the write

Enabling a source whose flag is already set drives the output a cycle
after the write, and the CPU sees it a cycle after that — as the comment
beside the delay constant already said. The constant was counted from
the CIA's caught-up cycle, which is the cycle *before* the write, so the
CPU saw the line a cycle early: the poll of the instruction right after
the write took the interrupt. Lorenz's `imr` checks exactly this pair
("irq in clock 2" must not happen, "no irq in clock 3" must not happen
either). The constant is now 3 from the caught-up cycle. `imr` passes.

## Disable: lands after timer B's output due in the write cycle

The mask write takes effect at the end of its cycle; an interrupt output
due in that same cycle is still driven with the mask as it was.
`cia-int-nmi`'s last column is that case for timer B. Timer B's pending
events are now brought to the write cycle before the mask changes (the
chip's own caught-up position is not moved). `cia-int-nmi` passes.

Timer A is deliberately left as it was: applying the same rule to it
makes `cia-icr-test2-continues/-oneshot` pass (a plain disabling store
in the output-due cycle: interrupt bit set) but breaks `dd0dtest` (test
11, `inc $dd0d,x`: read, dummy write, then the disabling write in the
output-due cycle: no NMI). Both programs were traced to the cycle and
their timer phases match our model, so the difference is real and lies
in what the read-modify-write's extra accesses do; no single rule I
tried satisfies both. A code comment records the tension;
`cia-icr-test2-*` keep failing as before.

## Not a bug: `irqnoack/ackraster`

It relies on the KERNAL's timer A having underflowed in the ~10 000
cycles between the program's `SEI` and its timer stop, which depends on
where in the timer's period the program starts. Shifting the harness's
start by 1–4 frames makes it time out as expected; at 0 or 5 frames it
fails. The program's own readme notes it could not be made to report a
pass. Left as is.

## No regressions

Every test-program suite re-run and compared program by program with the
previous branch head: VICII 274 and CIA 104 identical; Lorenz + banking
295 identical apart from `imr` (fail → pass); CPU + interrupts 142
identical apart from `cia-int-nmi` (fail → pass). Unit suites green. Two
new tests in `C64DeviceAccessTimingTests`: the enable seen by the second
poll after the write, and timer B's disable in the output-due cycle with
the negative case one cycle earlier. `C64ExecuteFrameBenchmark` at
parity (the changed code runs only on interrupt-control writes). Docs:
the CIA timing paragraph on the C64 libraries page updated.
…the VICE testbench run with the legacy pixel generator (#355)

A new documentation page, **C64 emulation accuracy and known
limitations** (`docs/systems/c64/accuracy.md`, in the navigation under
the C64 system after *Libraries*, and linked from the overview). It is
for readers who want to know, at the level of individual cycles, how
close the emulation is to the real machine and where it stops. Each
point names the VICE test programs that show it, so a reader can check
it themselves.

## What the page covers

- **Models:** 6569 (PAL) and 6567R8 (NTSC) VIC-II; the 6567R56A and
8565/8562 are not modelled. The original 6526 CIA; the 6526A ("new CIA")
is not, and cannot be selected.
- **CPU:** coverage of undocumented opcodes per compatibility profile;
the `ANE`/`LXA` magic constant; undriven 6510 port bits that hold their
value without fading; `JAM`.
- **VIC-II:** graphics memory written in the cycles it is fetched shows
a few cycles early (the picture is drawn after each instruction); the
sprite collision interrupt is raised at the end of the line rather than
at the pixel; sprite-gap cases; the 8565 light pen.
- **The legacy pixel generator:** what the faster generator does not
show — 8-pixel blocks, XSCROLL/mode bits/pointers read once per line,
sprites as whole bands, sprite-to-background collisions not taken from
the drawn graphics — plus the effect of turning per-line sprites off.
- **CIA:** not modelled — timer cascade, the CNT input, PB6/PB7 timer
output, the serial shift register, the time-of-day clock. Three known
timing differences: `loadth`, `flipos`, and a timer-A interrupt disabled
in the cycle its output goes active.
- Programs that differ for reasons not yet analysed
(`irqdma/test5`–`7b`, `cputiming`), and `ackraster`, which depends on
where it starts in the KERNAL timer's period.

## Testbench harness

A new `--pixel-generator sequencer|legacy` option (documented in the
development guide) runs the VICE programs with the legacy pixel
generator. The page's statement about the legacy generator is based on a
run with it: 126 of the 274 VIC-II programs pass with it against 226
with the default generator, and no program passes only with the legacy
one. The suites named on the page are the ones where the difference
lies.

## Checks

`mkdocs build --strict` passes, with every link and anchor on the new
page resolving. The only code change is the harness option; the harness
builds with no warnings. The claims about opcode coverage per profile,
the TOD and SDR registers, and how background collisions are computed
with the legacy generator were each checked against the code.
…he VIC-II timing rules (#356)

Rewords two older mentions that attributed VIC-II timing to VICE itself:

- `docs/systems/c64/libraries.md`: the points within a cycle at which
register changes reach the output are described as the chip's observed
behaviour as VICE's VICII test programs show it, whose reference
pictures the output is checked against.
- `Vic2BusStallTests.cs`: the 6567R8 sprite fetch cycles are stated as
chip behaviour, matching the comment in `Vic2BusStalls.cs`.

Documentation and comment only; no behaviour change. Strict docs build
passes; `Vic2BusStallTests` pass (32/32).
…ork in the sequencer pixel generator (#358)

Two parts of the sequencer pixel generator did per-pixel work that the
steady case does not need:

- **Scrolled blocks.** With XSCROLL set, every 8-pixel block was drawn
pixel by pixel. The block is now composed from the previous byte's
decoded colour codes and the new byte's when the modes and XSCROLL have
not changed since that byte was loaded; otherwise it is drawn pixel by
pixel as before.
- **Line colour resolve.** Only blocks marked as repeats were copied;
now any 8-pixel block whose codes equal the block before it is copied
instead of looked up. The repeat marks are removed.

**Output unchanged**
- Giana Sisters, 4,000 frames with scripted input: every frame hashes
the same as before, with both pixel generators.
- VICE VICII suite (274 programs, PAL and NTSC): same verdict for every
program (226 pass), byte-identical pictures.
- `Highbyte.DotNet6502.Systems.Tests`: 1,545 passed, 0 failed.

**Performance** (Giana Sisters, render and audio, Apple M1, ms per
frame; details in `RESULTS.md`)

| Build | Title | Game |
|-------|------:|-----:|
| master | 1.035 | 1.033 |
| before | 1.437 | 1.373 |
| after | 1.339 (−7%) | 1.322 (−4%) |
@sonarqubecloud

Copy link
Copy Markdown

@highbyte
highbyte merged commit 4abca4e into master Sep 28, 2026
10 checks passed
@highbyte
highbyte deleted the feature/cpu-cycle-engine branch September 28, 2026 14:39
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant