Charge the memory controller for the memory it moves

A blit cost ten cycles, which were the five port writes that set it up. The
quarter of a kilobyte that moved cost nothing, and no hardware moves a quarter
of a kilobyte for nothing.

BANKS ARE SEPARATE MEMORIES, AND THAT IS WHAT SETS THE RATE. A move between two
of them can overlap its read and its write - fetch the next byte while the last
one is stored - so it settles at a byte a cycle. A move within one bank cannot,
and costs two. A fill has nothing to read and costs one whatever the banks are.
The odd cycle on each is the pipeline filling.

That is not a modelling choice so much as a reading of the structure the machine
already has: a Program to Data blit is inherently twice the rate of a Data to
Data one, and it is legible why.

Measured: 256 bytes is 297 cycles across banks and 518 within one, both
including the instructions that ask for it.

WHAT IT TAUGHT, which was not what I expected. Charging for movement costs the
native assembler 0.4 per cent and costs directory work 13.4. The assembler reads
a block and then thinks about it for a long time, so the move is amortised into
nothing; the filesystem reads a block in order to look at it and does nothing
else in between.

So the case for a blitter that runs alongside the CPU is weaker than it sounds.
Concurrency pays when there is other work to do during the transfer, and the
place that spends its time moving memory is exactly the place with nothing else
to do - it blits a block precisely so that it can read it. What that workload
wants is a FASTER controller, not a concurrent one: a wider data path halves the
wait, and the machine is waiting either way.

Video is the case that would still want concurrency, since a frame can be moved
while the next one is worked out. That is an argument about software nobody has
written yet, and it is now an argument with numbers on the other side of it.

The byte at a time port is charged too, for the byte it moves beyond reaching
the port. Nothing polls CTRL_STATUS, so the transfer stalls whoever asked for
it, which is the conservative reading and the one the software already assumes.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01E2JrLzFvuFX9fgi1LDRjrW
This commit is contained in:
Anachronaut
2026-08-25 21:11:47 -04:00
co-authored by Claude Opus 5
parent f1e5cc46f6
commit e0cf0a9a25
4 changed files with 53 additions and 2 deletions
+12 -1
View File
@@ -102,7 +102,18 @@ anybody could build - and it is the emulator's job to be the thing the hardware
against.
The average SplitBit instruction costs 3.72 cycles, measured over the native assembler
assembling a program. Whether real hardware would overlap a fetch with the end of the
assembling a program.
**The memory controller is charged for what it moves**, on the same terms. Banks are
separate memories, and that is what sets the rate: a move between two of them can overlap
its read and its write, so it settles at a byte a cycle, while a move within one bank cannot
and costs two. A fill has nothing to read and costs one. So a 256 byte block is 257 cycles
between banks and 513 within one, against the ten it used to cost - which was the five port
writes that set it up and nothing for the quarter of a kilobyte that moved.
The transfer stalls the program that asked for it. Whether hardware would let the two run at
once is left open, the same way pipelining is: the memories are separate, so it plausibly
could, and the measurements say it would buy less than it sounds like. Whether real hardware would overlap a fetch with the end of the
previous instruction is left open, and deliberately: this is the conservative model, and
pipelining is a decision to make while drawing the hardware rather than one to inherit from
an emulator.