Repository navigation
Conversation
danbugs
requested review from
andreiltd,
dblnz,
devigned,
jprendes,
jsturtevant,
ludfjig,
simongdavies,
squillace and
syntactically
as code owners
October 7, 2026 23:57
Contributor
There was a problem hiding this comment.
🔵 Needs a closer look
The pagemap-driven reset policy affects sandbox isolation and needs fixes plus human validation.
3 open findings
What changed in this PR
This experiment targets faster KVM snapshot restores by retaining and zeroing selected scratch pages in Hyperlight’s host memory manager.
Changes:
- Adaptive pagemap-based scratch resets with platform fallbacks.
- Per-sandbox reset state, memory-behavior tests, and changelog documentation.
| File | Description |
|---|---|
| src/hyperlight_host/src/mem/shared_mem.rs | Selects adaptive KVM resets with zeroing fallbacks. |
| src/hyperlight_host/src/mem/scratch_reset.rs | Implements residency tracking, reset policies, and tests. |
| src/hyperlight_host/src/mem/mod.rs | Provides platform-gated reset modules. |
| src/hyperlight_host/src/mem/mgr.rs | Maintains reset state across scratch restores. |
| CHANGELOG.md | Documents retained KVM scratch memory. |
🧠 Review effort: Balanced
Give feedback about Copilot approvals in this survey to enter a drawing for a $150 gift card.
danbugs
force-pushed
the
resident-scratch-reset
branch
from
October 8, 2026 00:11
74aa786 to
b907442
Compare
This comment has been minimized.
This comment has been minimized.
danbugs
force-pushed
the
resident-scratch-reset
branch
from
October 8, 2026 20:40
d278d83 to
592e155
Compare
This comment has been minimized.
This comment has been minimized.
This comment has been minimized.
This comment has been minimized.
danbugs
force-pushed
the
resident-scratch-reset
branch
2 times, most recently
from
October 9, 2026 00:37
7a2bee3 to
cb8bbaf
Compare
This comment has been minimized.
This comment has been minimized.
A restore zeroed scratch by dropping it with MADV_DONTNEED on KVM (fill(0) in builds with mshv3), by fill(0) on MSHV, and by mapping fresh memory on Windows (#1765). Each costs time in the size of scratch, or a fault on every page the next run touches. On KVM, every reset now reads /proc/self/pagemap for all of scratch. A page with memory of its own that holds data is zeroed and kept, so the next run takes no faults on it. One that holds only zeros was not written since the last reset and is dropped, so what stays resident follows what runs write. Swapped pages and shared pages holding data are dropped. The rule is the same from the first restore on. Scratch is backed 4 KiB at a time on KVM, so a reset zeroes only what was touched. On Windows, scratch of up to 16 MiB is now zeroed in place, which is about 2x faster than a fresh mapping (over 5x when the guest writes a MiB or more); larger scratch is still replaced, so it does not all become resident. HostSharedMemory now notes the pages its writes touch, and zero_written zeroes only a given set of pages plus those, for resets that know what the guest wrote (next commits). Signed-off-by: danbugs <danilochiarlone@gmail.com>
Add a DirtyLog supertrait of VirtualMachine: how a VM tracks the pages the guest writes, and reading and clearing that log for a range. On MSHV (x86_64), tracking is a partition property switched on and off. Reads use get_dirty_log with clear; stopping sets the bits back first, which MSHV requires. On WHP (x86_64 and ARM64), scratch is mapped with WHvMapGpaRangeFlagTrackDirtyPages and read with WHvQueryGpaRangeDirtyBitmap, which clears what it reads. If the tracked mapping fails, scratch is mapped as before and tracking is not tried again for the VM. Other backends do not track. A backend that claims to but does not read reports every page, so a reset zeroes too much, never too little. Signed-off-by: danbugs <danilochiarlone@gmail.com>
Where the hypervisor logs the guest's writes, a restore now zeroes only the pages written since the last restore: those in the log, the pages host writes touched, and the page-table copy (#1766). Tracking starts when scratch is mapped, before the guest runs, so the first restore uses it too (hyperlight_vm/dirty_log.rs): - WHP: always. Tracking costs the guest nothing measurable there. The log is cleared when scratch is mapped, so only written pages are ever touched and the rest of scratch never becomes resident. - MSHV: tracking makes the guest's first write to each page after a read fault to the hypervisor (about 1 us nested), and reading the log costs time in the size of scratch. Measured against zeroing all of scratch, that loses below 2 MiB of scratch, and once a run writes more than about a tenth of scratch (a quarter from 32 MiB, where zeroing gets about 3x slower per MiB). So scratch under 2 MiB is not tracked, and tracking stops after such a run until scratch is mapped again. The first log after mapping reports every page, so the guest sets up without faults. - KVM keeps its in-place reset; no log is read. Signed-off-by: danbugs <danilochiarlone@gmail.com>
The restore benchmarks use 352 KiB to 1 MiB of scratch, where resetting it costs little however it is done. Add call_with_restore with 64 and 256 MiB of scratch, with a guest that writes a few pages (Echo) and one that writes 1 MiB a call, and report them, with the default size, in pull request comments. Signed-off-by: danbugs <danilochiarlone@gmail.com>
danbugs
force-pushed
the
resident-scratch-reset
branch
from
October 11, 2026 00:06
e4d8f03 to
e53c815
Compare
danbugs
added a commit
to hyperlight-dev/hyperlight-unikraft
that referenced
this pull request
Oct 11, 2026
Patch hyperlight-host and hyperlight-common to danbugs/hyperlight c854bb43 (0.17.0 plus the scratch reset of hyperlight-dev/hyperlight#1901), to measure it in CI. Signed-off-by: danbugs <danilochiarlone@gmail.com>
Benchmark ResultsMeasured commit: kvm / amd (Linux) (➖ stable)No benchmark improved or regressed. Benchmark Resultsfunction_call_codec
guest_calls
payload_allocation
sandboxes
slot_pool
snapshot_files
virtq_readonly
virtq_readwrite
kvm / intel (Linux) (❌ 2.00x)Top regressions
Benchmark Resultsfunction_call_codec
guest_calls
payload_allocation
sandboxes
slot_pool
snapshot_files
virtq_readonly
virtq_readwrite
mshv3 / amd (Linux) (➖ stable)No benchmark improved or regressed. Benchmark Resultsfunction_call_codec
guest_calls
payload_allocation
sandboxes
slot_pool
snapshot_files
virtq_readonly
virtq_readwrite
mshv3 / intel (Linux) (➖ stable)No benchmark improved or regressed. Benchmark Resultsfunction_call_codec
guest_calls
payload_allocation
sandboxes
slot_pool
snapshot_files
virtq_readonly
virtq_readwrite
hyperv-ws2025 / amd (Windows) (🌟 5.47x)Top improvements
Benchmark Resultsfunction_call_codec
guest_calls
payload_allocation
sandboxes
slot_pool
snapshot_files
virtq_readonly
virtq_readwrite
hyperv-ws2025 / intel (Windows) (🌟 6.00x)Top improvements
Benchmark Resultsfunction_call_codec
guest_calls
payload_allocation
sandboxes
slot_pool
snapshot_files
virtq_readonly
virtq_readwrite
Reported by |
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.


No description provided.