Commits · 7f582ab6d8116ce8db5792c219a278519deae6ad · nexedi / linux

10 Sep, 2009 40 commits

KVM: VMX: Avoid to return ENOTSUPP to userland · 7f582ab6

Jan Kiszka authored Jul 22, 2009

Choose some allowed error values for the cases VMX returned ENOTSUPP so
far as these values could be returned by the KVM_RUN IOCTL.
Signed-off-by: Jan Kiszka <jan.kiszka@siemens.com>
Signed-off-by: Marcelo Tosatti <mtosatti@redhat.com>

7f582ab6

KVM: Drop obsolete cpu_get/put in make_all_cpus_request · e601e3be

Jan Kiszka authored Jul 20, 2009

spin_lock disables preemption, so we can simply read the current cpu.
Signed-off-by: Jan Kiszka <jan.kiszka@siemens.com>
Signed-off-by: Marcelo Tosatti <mtosatti@redhat.com>

e601e3be

KVM: PIT: Unregister ack notifier callback when freeing · 84fde248
Gleb Natapov authored Jul 16, 2009
```
Signed-off-by: Gleb Natapov <gleb@redhat.com>
Signed-off-by: Marcelo Tosatti <mtosatti@redhat.com>
```
84fde248

KVM: VMX: Introduce KVM_SET_IDENTITY_MAP_ADDR ioctl · b927a3ce

Sheng Yang authored Jul 21, 2009

Now KVM allow guest to modify guest's physical address of EPT's identity mapping page.

(change from v1, discard unnecessary check, change ioctl to accept parameter
address rather than value)
Signed-off-by: Sheng Yang <sheng@linux.intel.com>
Signed-off-by: Marcelo Tosatti <mtosatti@redhat.com>

b927a3ce

KVM: x86: use kvm_get_gdt() and kvm_read_ldt() · b792c344

Akinobu Mita authored Jul 19, 2009

Use kvm_get_gdt() and kvm_read_ldt() to reduce inline assembly code.

Cc: Avi Kivity <avi@redhat.com>
Cc: kvm@vger.kernel.org
Signed-off-by: Akinobu Mita <akinobu.mita@gmail.com>
Signed-off-by: Marcelo Tosatti <mtosatti@redhat.com>

b792c344

KVM: x86: use get_desc_base() and get_desc_limit() · 46a359e7

Akinobu Mita authored Jul 18, 2009

Use get_desc_base() and get_desc_limit() to get the base address and
limit in desc_struct.

Cc: Avi Kivity <avi@redhat.com>
Cc: kvm@vger.kernel.org
Signed-off-by: Akinobu Mita <akinobu.mita@gmail.com>
Signed-off-by: Marcelo Tosatti <mtosatti@redhat.com>

46a359e7

KVM: s390: remove unused structs · decde80b

Gleb Natapov authored Jul 12, 2009

They are not used by common code without defines which s390 does not
have.
Signed-off-by: Gleb Natapov <gleb@redhat.com>
Signed-off-by: Marcelo Tosatti <mtosatti@redhat.com>

decde80b

KVM: MMU: fix missing locking in alloc_mmu_pages · 6a1ac771

Marcelo Tosatti authored Jul 15, 2009

n_requested_mmu_pages/n_free_mmu_pages are used by
kvm_mmu_change_mmu_pages to calculate the number of pages to zap.

alloc_mmu_pages, called from the vcpu initialization path, modifies this
variables without proper locking, which can result in a negative value
in kvm_mmu_change_mmu_pages (say, with cpu hotplug).
Signed-off-by: Marcelo Tosatti <mtosatti@redhat.com>

6a1ac771

KVM: Discard unnecessary kvm_mmu_flush_tlb() in kvm_mmu_load() · 3662cb1c

Sheng Yang authored Jul 09, 2009

set_cr3() should already cover the TLB flushing.
Signed-off-by: Sheng Yang <sheng@linux.intel.com>
Signed-off-by: Marcelo Tosatti <mtosatti@redhat.com>

3662cb1c

KVM: silence lapic kernel messages that can be triggered by a guest · 4088bb3c

Gleb Natapov authored Jul 08, 2009

Some Linux versions (f8) try to read EOI register that is write only.
Signed-off-by: Gleb Natapov <gleb@redhat.com>
Signed-off-by: Marcelo Tosatti <mtosatti@redhat.com>

4088bb3c

KVM: Reduce runnability interface with arch support code · a1b37100

Gleb Natapov authored Jul 09, 2009

Remove kvm_cpu_has_interrupt() and kvm_arch_interrupt_allowed() from
interface between general code and arch code. kvm_arch_vcpu_runnable()
checks for interrupts instead.
Signed-off-by: Gleb Natapov <gleb@redhat.com>
Signed-off-by: Avi Kivity <avi@redhat.com>

a1b37100

KVM: Move kvm_cpu_get_interrupt() declaration to x86 code · 0b71785d

Gleb Natapov authored Jul 09, 2009

It is implemented only by x86.
Signed-off-by: Gleb Natapov <gleb@redhat.com>
Signed-off-by: Avi Kivity <avi@redhat.com>

0b71785d

KVM: Move exception handling to the same place as other events · b59bb7bd
Gleb Natapov authored Jul 09, 2009
```
Signed-off-by: Gleb Natapov <gleb@redhat.com>
Signed-off-by: Avi Kivity <avi@redhat.com>
```
b59bb7bd

KVM: MMU: Fix MMU_DEBUG compile breakage · a205bc19

Joerg Roedel authored Jul 09, 2009

Signed-off-by: Joerg Roedel <joerg.roedel@amd.com>
Signed-off-by: Avi Kivity <avi@redhat.com>

a205bc19

KVM: add ioeventfd support · d34e6b17

Gregory Haskins authored Jul 07, 2009

ioeventfd is a mechanism to register PIO/MMIO regions to trigger an eventfd
signal when written to by a guest.  Host userspace can register any
arbitrary IO address with a corresponding eventfd and then pass the eventfd
to a specific end-point of interest for handling.

Normal IO requires a blocking round-trip since the operation may cause
side-effects in the emulated model or may return data to the caller.
Therefore, an IO in KVM traps from the guest to the host, causes a VMX/SVM
"heavy-weight" exit back to userspace, and is ultimately serviced by qemu's
device model synchronously before returning control back to the vcpu.

However, there is a subclass of IO which acts purely as a trigger for
other IO (such as to kick off an out-of-band DMA request, etc).  For these
patterns, the synchronous call is particularly expensive since we really
only want to simply get our notification transmitted asychronously and
return as quickly as possible.  All the sychronous infrastructure to ensure
proper data-dependencies are met in the normal IO case are just unecessary
overhead for signalling.  This adds additional computational load on the
system, as well as latency to the signalling path.

Therefore, we provide a mechanism for registration of an in-kernel trigger
point that allows the VCPU to only require a very brief, lightweight
exit just long enough to signal an eventfd.  This also means that any
clients compatible with the eventfd interface (which includes userspace
and kernelspace equally well) can now register to be notified. The end
result should be a more flexible and higher performance notification API
for the backend KVM hypervisor and perhipheral components.

To test this theory, we built a test-harness called "doorbell".  This
module has a function called "doorbell_ring()" which simply increments a
counter for each time the doorbell is signaled.  It supports signalling
from either an eventfd, or an ioctl().

We then wired up two paths to the doorbell: One via QEMU via a registered
io region and through the doorbell ioctl().  The other is direct via
ioeventfd.

You can download this test harness here:

ftp://ftp.novell.com/dev/ghaskins/doorbell.tar.bz2

The measured results are as follows:

qemu-mmio:       110000 iops, 9.09us rtt
ioeventfd-mmio: 200100 iops, 5.00us rtt
ioeventfd-pio:  367300 iops, 2.72us rtt

I didn't measure qemu-pio, because I have to figure out how to register a
PIO region with qemu's device model, and I got lazy.  However, for now we
can extrapolate based on the data from the NULLIO runs of +2.56us for MMIO,
and -350ns for HC, we get:

qemu-pio:      153139 iops, 6.53us rtt
ioeventfd-hc: 412585 iops, 2.37us rtt

these are just for fun, for now, until I can gather more data.

Here is a graph for your convenience:

http://developer.novell.com/wiki/images/7/76/Iofd-chart.png

The conclusion to draw is that we save about 4us by skipping the userspace
hop.

--------------------
Signed-off-by: Gregory Haskins <ghaskins@novell.com>
Acked-by: Michael S. Tsirkin <mst@redhat.com>
Signed-off-by: Avi Kivity <avi@redhat.com>

d34e6b17

KVM: make io_bus interface more robust · 090b7aff

Gregory Haskins authored Jul 07, 2009

Today kvm_io_bus_regsiter_dev() returns void and will internally BUG_ON
if it fails.  We want to create dynamic MMIO/PIO entries driven from
userspace later in the series, so we need to enhance the code to be more
robust with the following changes:

   1) Add a return value to the registration function
   2) Fix up all the callsites to check the return code, handle any
      failures, and percolate the error up to the caller.
   3) Add an unregister function that collapses holes in the array
Signed-off-by: Gregory Haskins <ghaskins@novell.com>
Acked-by: Michael S. Tsirkin <mst@redhat.com>
Signed-off-by: Avi Kivity <avi@redhat.com>

090b7aff

KVM: add module parameters documentation · fef07aae

Andre Przywara authored Jul 10, 2009

Signed-off-by: Andre Przywara <andre.przywara@amd.com>
Signed-off-by: Avi Kivity <avi@redhat.com>

fef07aae

KVM: PIT support for HPET legacy mode · e9f42757

Beth Kon authored Jul 07, 2009

When kvm is in hpet_legacy_mode, the hpet is providing the timer
interrupt and the pit should not be. So in legacy mode, the pit timer
is destroyed, but the *state* of the pit is maintained. So if kvm or
the guest tries to modify the state of the pit, this modification is
accepted, *except* that the timer isn't actually started. When we exit
hpet_legacy_mode, the current state of the pit (which is up to date
since we've been accepting modifications) is used to restart the pit
timer.

The saved_mode code in kvm_pit_load_count temporarily changes mode to
0xff in order to destroy the timer, but then restores the actual
value, again maintaining "current" state of the pit for possible later
reenablement.

[avi: add some reserved storage in the ioctl; make SET_PIT2 IOW]
[marcelo: fix memory corruption due to reserved storage]
Signed-off-by: Beth Kon <eak@us.ibm.com>
Signed-off-by: Marcelo Tosatti <mtosatti@redhat.com>
Signed-off-by: Avi Kivity <avi@redhat.com>

e9f42757

KVM: Always report x2apic as supported feature · 0d1de2d9

Gleb Natapov authored Jul 12, 2009

We emulate x2apic in software, so host support is not required.
Signed-off-by: Gleb Natapov <gleb@redhat.com>
Signed-off-by: Avi Kivity <avi@redhat.com>

0d1de2d9

KVM: No need to kick cpu if not in a guest mode · c7f0f24b

Gleb Natapov authored Jul 07, 2009

This will save a couple of IPIs.
Signed-off-by: Gleb Natapov <gleb@redhat.com>
Acked-by: Marcelo Tosatti <mtosatti@redhat.com>
Signed-off-by: Avi Kivity <avi@redhat.com>

c7f0f24b

KVM: Add trace points in irqchip code · 1000ff8d

Gleb Natapov authored Jul 07, 2009

Add tracepoint in msi/ioapic/pic set_irq() functions,
in IPI sending and in the point where IRQ is placed into
apic's IRR.
Signed-off-by: Gleb Natapov <gleb@redhat.com>
Signed-off-by: Avi Kivity <avi@redhat.com>

1000ff8d

KVM: ignore msi request if !level · 07fb8bb2

Michael S. Tsirkin authored Jul 05, 2009

Irqfd sets level for interrupt to 1 and then to 0.
For MSI, check level so that a single message is sent.
Signed-off-by: Michael S. Tsirkin <mst@redhat.com>
Signed-off-by: Avi Kivity <avi@redhat.com>

07fb8bb2

KVM: fix MMIO_CONF_BASE MSR access · f7c6d140

Andre Przywara authored Jul 02, 2009

Some Windows versions check whether the BIOS has setup MMI/O for
config space accesses on AMD Fam10h CPUs, we say "no" by returning 0 on
reads and only allow disabling of MMI/O CfgSpace setup by igoring "0" writes.
Signed-off-by: Andre Przywara <andre.przywara@amd.com>
Signed-off-by: Avi Kivity <avi@redhat.com>

f7c6d140

KVM: Trace shadow page lifecycle · f691fe1d
Avi Kivity authored Jul 06, 2009
```
Create, sync, unsync, zap.
Signed-off-by: Avi Kivity <avi@redhat.com>
```
f691fe1d

KVM: Document basic API · 9c1b96e3

Avi Kivity authored Jun 09, 2009

Document the basic API corresponding to the 2.6.22 release.
Signed-off-by: Avi Kivity <avi@redhat.com>

9c1b96e3

KVM: MMU: Trace guest pagetable walker · 07420171
Avi Kivity authored Jul 06, 2009
```
Signed-off-by: Avi Kivity <avi@redhat.com>
```
07420171

Revert "KVM: x86: check for cr3 validity in ioctl_set_sregs" · dc7e795e

Jan Kiszka authored Jul 01, 2009

This reverts commit 6c20e1442bb1c62914bb85b7f4a38973d2a423ba.

To my understanding, it became obsolete with the advent of the more
robust check in mmu_alloc_roots (89da4ff17f). Moreover, it prevents
the conceptually safe pattern

 1. set sregs
 2. register mem-slots
 3. run vcpu

by setting a sticky triple fault during step 1.
Signed-off-by: Jan Kiszka <jan.kiszka@siemens.com>
Signed-off-by: Avi Kivity <avi@redhat.com>

dc7e795e

KVM: handle AMD microcode MSR · 6098ca93

Andre Przywara authored Jul 03, 2009

Windows 7 tries to update the CPU's microcode on some processors,
so we ignore the MSR write here. The patchlevel register is already handled
(returning 0), because the MSR number is the same as Intel's.
Signed-off-by: Andre Przywara <andre.przywara@amd.com>
Signed-off-by: Avi Kivity <avi@redhat.com>

6098ca93

KVM: Fix apic_mmio_write return for unaligned write · 756975bb

Sheng Yang authored Jul 06, 2009

Some in-famous OS do unaligned writing for APIC MMIO, and the return value
has been missed in recent change, then the OS hangs.
Signed-off-by: Sheng Yang <sheng@linux.intel.com>
Signed-off-by: Avi Kivity <avi@redhat.com>

756975bb

KVM: Use temporary variable to shorten lines. · 70f93dae

Gleb Natapov authored Jul 05, 2009

Cosmetic only. No logic is changed by this patch.
Signed-off-by: Gleb Natapov <gleb@redhat.com>
Signed-off-by: Avi Kivity <avi@redhat.com>

70f93dae

KVM: x2apic interface to lapic · 0105d1a5

Gleb Natapov authored Jul 05, 2009

This patch implements MSR interface to local apic as defines by x2apic
Intel specification.
Signed-off-by: Gleb Natapov <gleb@redhat.com>
Signed-off-by: Avi Kivity <avi@redhat.com>

0105d1a5

KVM: Add Directed EOI support to APIC emulation · fc61b800

Gleb Natapov authored Jul 05, 2009

Directed EOI is specified by x2APIC, but is available even when lapic is
in xAPIC mode.
Signed-off-by: Gleb Natapov <gleb@redhat.com>
Signed-off-by: Avi Kivity <avi@redhat.com>

fc61b800

KVM: Trace apic registers using their symbolic names · cb247721
Avi Kivity authored Jul 01, 2009
```
Signed-off-by: Avi Kivity <avi@redhat.com>
```
cb247721
KVM: Trace mmio · aec51dc4
Avi Kivity authored Jul 01, 2009
```
Signed-off-by: Avi Kivity <avi@redhat.com>
```
aec51dc4

KVM: Ignore PCI ECS I/O enablement · c323c0e5

Andre Przywara authored Jun 24, 2009

Linux guests will try to enable access to the extended PCI config space
via the I/O ports 0xCF8/0xCFC on AMD Fam10h CPU. Since we (currently?)
don't use ECS, simply ignore write and read attempts.
Signed-off-by: Andre Przywara <andre.przywara@amd.com>
Signed-off-by: Avi Kivity <avi@redhat.com>

c323c0e5

KVM: Trace irq level and source id · ae8c1c40
Avi Kivity authored Jul 01, 2009
```
Signed-off-by: Avi Kivity <avi@redhat.com>
```
ae8c1c40

KVM: fix lock imbalance · 27c4ba60

Jiri Slaby authored Jun 29, 2009

There is a missing unlock on one fail path in ioapic_mmio_write,
fix that.
Signed-off-by: Jiri Slaby <jirislaby@gmail.com>
Signed-off-by: Avi Kivity <avi@redhat.com>

27c4ba60

KVM: document lock nesting rule · 22fc0294

Michael S. Tsirkin authored Jun 29, 2009

Document kvm->lock nesting within kvm->slots_lock
Signed-off-by: Michael S. Tsirkin <mst@redhat.com>
Signed-off-by: Avi Kivity <avi@redhat.com>

22fc0294

KVM: remove in_range from io devices · bda9020e

Michael S. Tsirkin authored Jun 29, 2009

This changes bus accesses to use high-level kvm_io_bus_read/kvm_io_bus_write
functions. in_range now becomes unused so it is removed from device ops in
favor of read/write callbacks performing range checks internally.

This allows aliasing (mostly for in-kernel virtio), as well as better error
handling by making it possible to pass errors up to userspace.
Signed-off-by: Michael S. Tsirkin <mst@redhat.com>
Signed-off-by: Avi Kivity <avi@redhat.com>

bda9020e

KVM: convert bus to slots_lock · 6c474694

Michael S. Tsirkin authored Jun 29, 2009

Use slots_lock to protect device list on the bus.  slots_lock is already
taken for read everywhere, so we only need to take it for write when
registering devices.  This is in preparation to removing in_range and
kvm->lock around it.
Signed-off-by: Michael S. Tsirkin <mst@redhat.com>
Signed-off-by: Avi Kivity <avi@redhat.com>

6c474694