|
[Date Prev][Date Next][Thread Prev][Thread Next][Date Index][Thread Index] Re: [PATCH v7 17/36] KVM: x86: Allow KVM master clock mode when TSCs are offset from each other
On Tue, 2026-08-11 at 11:41 -0700, Sean Christopherson wrote: > OMG, I hate this code. After literally hours of staring at this, and even > typing > up a lengthy example of why guest time would go off the rails, I finally > spotted > that l1_tsc_offset is accounted for by the call to kvm_read_l1_tsc(). FML. > > Thanks for being patient and not flaming me too much :-) Haha, no judgement. We *all* hate this code. That's why I threw my toys out of the pram and went on this crusade to clean it up a bit. > > > > And your variant just added a dependency on wallclock time back into it > > > > > > Can you elaborate? I'm guessing I don't entirely understand what you > > > mean by > > > wallclock time. > > > > The system_time field? The unspecified might-be-UTC-might-have-leap-seconds > > one :) > > Ok, I think I finally understand the goal. I got turned around by the > combination > of the name SET_CLOCK_GUEST and the full pvclock structure being passed to the > guest. I was expecting SET_CLOCK_GUEST to literally set the entire clock, > e.g. > mul+shift, timestamp, etc. That's an implementation detail. It is literally getting the clock as the guest sees it, and setting it again on the destination from the same guest-ABI pvclock structure. From the userspace point of view those actions *are* symmetrical. I'd actually *like* it to just be a memcpy at both ends, even on the SET side, just copying what userspace provides into what we offer to the guest as its pvclock. But as well as wanting to do some sanity checking, we also live in a world where we might have to switch to the non-masterclock mode at any time, and we have to ingest the information into the per-VM kvmclock setup in a way that the kernel "understands", and that's why it ends up implemented the way it is. I looked at rewriting the masterclock base information from what userspace is providing, but there are *host* TSC values in there, and it ended up in some cases wanting to set ka->master_cycle_now to a value which is *negative* on the new host, and I didn't want to exercise that wrap-around path. So instead we just adjust ka->kvmclock_offset to give appropriate results based on the existing masterclock base. Each vCPU's pvclock is then *regenerated* from the VM-side kvmclock data, giving rise to that annoying ±1ns discrepancy that I whined about a while back, but didn't give in to my perfectionism and eliminate... yet. As far as userspace is concerned, KVM_SET_CLOCK_GUEST *does* set the entire clock (at least the relationship between guest TSC and kvmclock nanoseconds, which is what it's for). It's just that the kernel then "tweaks" it a little bit to give a slightly different y=m(x-x')+c equation which is still within the noise of the original. And the kernel actually does that kind of 'tweak' all the time. Although we're working on narrowing them down because "within the noise" is in the eye of the beholder; Dongli had some patches for that which I think I rounded up and included? > But all of that metadata is just a means to an end: the one and only goal is > to > calculate the per-VM kvmclock_offset for the "new" host's TSC+time snapshot, > by > computing the nanoseconds delta for the new snapshot as if it the guest > observed > the TSC while running on the old host. That is currently how it is implemented. It isn't the API contract. > And that is done in the kernel instead of in userspace to minimize the amount > of > slop introduced due to delay between taking the snapshot and computing the > offset. Huh? There should be no delays here. If *anything* in this new code is done with something other than a *simultaneous* reading of TSC and ktime that I sweated blood and tears and got shouted at by Thomas for, then that *is* something I care about... Attachment:
smime.p7s
|
![]() |
Lists.xenproject.org is hosted with RackSpace, monitoring our |