Commits · 1090c056c5eb6d5335cceb381683e77ac24c71ab · Kirill Smelkov / linux

14 Oct, 2010 40 commits

drbd: drbd_md_sync before calling user space helpers · 1090c056

Lars Ellenberg authored Jul 19, 2010

Just in case we have some pending meta data changes to sync, do it
before we call our userland helper, as that may take some time,
or even cause a hard reboot.
Signed-off-by: Philipp Reisner <philipp.reisner@linbit.com>
Signed-off-by: Lars Ellenberg <lars.ellenberg@linbit.com>

1090c056

drbd: fix race on meta-data update, addendum · ee15b038

Lars Ellenberg authored Sep 03, 2010

addendum to baa33ae4eaa4477b60af7c434c0ddd1d182c1ae7

The race:
    drbd_md_sync()
	if (!test_and_clear_bit(MD_DIRTY, &mdev->flags))
		return;
    ==> RACE with drbd_md_mark_dirty() rearming the timer.
	del_timer(&mdev->md_sync_timer);

    Fixed by moving the del_timer before the test_and_clear_bit.

Additionally only rearm the timer in drbd_md_mark_dirty, if MD_DIRTY was
not already set, reduce the grace period from five to one second, and
add an ifdef'ed debuging aid to find code paths missing an explicit
drbd_md_sync, if any, as those are the only relevant ones for this race.
Signed-off-by: Philipp Reisner <philipp.reisner@linbit.com>
Signed-off-by: Lars Ellenberg <lars.ellenberg@linbit.com>

ee15b038

drbd: Removed a race that could cause unexpected execution of w_make_resync_request() · 63106d3c

Philipp Reisner authored Sep 01, 2010

The actual race happened int the drbd_start_resync() function. Where
drbd_resync_finished() -> __drbd_set_state() set STOP_SYNC_TIMER and
armed the timer.

If the timer fired before execution reaches the mod_timer statement
at the end of drbd_start_resync() the latter would cause an
unexpected call to w_make_resync_request().

Removed the STOP_SYNC_TIMER bit, and base it on the connection state.

The STOP_SYNC_TIMER bit probably originates probably the time before
the state engine.
Signed-off-by: Philipp Reisner <philipp.reisner@linbit.com>
Signed-off-by: Lars Ellenberg <lars.ellenberg@linbit.com>

63106d3c

drbd: implicitly create unconfigured devices on sync-after dependencies · ef50a3e3

Lars Ellenberg authored Sep 01, 2010

If pacemaker (for example) decided to initialize minor devices not in
the exact sync-after dependency order, the configuration partially
failed with an error "The sync-after minor number is invalid". (Bugz. #322)

We can avoid that by implicitly creating unconfigured minor devices,
if others depend on them.
Signed-off-by: Philipp Reisner <philipp.reisner@linbit.com>
Signed-off-by: Lars Ellenberg <lars.ellenberg@linbit.com>

ef50a3e3

drbd: fix race on meta-data update · 3f3a9b84

Lars Ellenberg authored Sep 01, 2010

The race:
	drbd_md_mark_dirty()
	drbd_md_sync()
		if (!test_and_clear_bit(MD_DIRTY, &mdev->flags))
			return;
		drbd_md_sync_page_io(mdev, mdev->ldev, sector, WRITE)
  ==> RACE
		clear_bit(MD_DIRTY, &mdev->flags); <== spurious

Fixed by removing the spurious clear_bit.
Signed-off-by: Philipp Reisner <philipp.reisner@linbit.com>
Signed-off-by: Lars Ellenberg <lars.ellenberg@linbit.com>

3f3a9b84

drbd: fix race between deconfiguring and reconfiguring network · c518d04f

Lars Ellenberg authored Sep 01, 2010

If a drbd_nl_net_conf hits the small window between the state change
to C_STANDALONE and the corresponding cleanup in after_state_ch,
that cleanup would throw away stuff we now need again,
and later trigger BUG_ON()s.

Fixed by properly serializing the new config request with
any pending cleanup.
Signed-off-by: Philipp Reisner <philipp.reisner@linbit.com>
Signed-off-by: Lars Ellenberg <lars.ellenberg@linbit.com>

c518d04f

drbd: Disable activity log updates when the whole device is out of sync · 0778286a

Philipp Reisner authored Aug 31, 2010

When the complete device is marked as out of sync, we can disable
updates of the on disk AL. Currently AL updates are only disabled
if one uses the "invalidate-remote" command on an unconnected,
primary device, or when at attach time all bits in the bitmap are
set.

As of now, AL updated do not get disabled when a all bits becomes
set due to application writes to an unconnected DRBD device.
While this is a missing feature, it is not considered important,
and might get added later.

BTW, after initializing a "one legged" DRBD device
drbdadm create-md resX
drbdadm -- --force primary resX
AL updates also get disabled, until the first connect.
Signed-off-by: Philipp Reisner <philipp.reisner@linbit.com>
Signed-off-by: Lars Ellenberg <lars.ellenberg@linbit.com>

0778286a

drbd: Actually allow BIOs up to 128k (was 32k). · d5373389

Philipp Reisner authored Aug 23, 2010

Now we have multiple BIOs per ee, packets with a 32 bit length field,
it gets time to use these goodies.
Signed-off-by: Philipp Reisner <philipp.reisner@linbit.com>
Signed-off-by: Lars Ellenberg <lars.ellenberg@linbit.com>

d5373389

drbd: receiving of big packets, for payloads between 64kByte and 4GByte · 02918be2
Philipp Reisner authored Aug 20, 2010
```
Signed-off-by: Philipp Reisner <philipp.reisner@linbit.com>
Signed-off-by: Lars Ellenberg <lars.ellenberg@linbit.com>
```
02918be2

drbd: Sending of big packets, for payloads from 64KByte to 4GByte · 0b70a13d

Philipp Reisner authored Aug 20, 2010

Signed-off-by: Philipp Reisner <philipp.reisner@linbit.com>
Signed-off-by: Lars Ellenberg <lars.ellenberg@linbit.com>

0b70a13d

drbd: Bugfix for regression introduced with f9bc8913c06022e · 204bba99

Philipp Reisner authored Aug 23, 2010

If we intent to use the block_id member of an epoch entry,
we may not use the digest member.
Signed-off-by: Philipp Reisner <philipp.reisner@linbit.com>
Signed-off-by: Lars Ellenberg <lars.ellenberg@linbit.com>

204bba99

drbd: Microfix: Assigning sector once is sufficient · 48acf868

Philipp Reisner authored Aug 23, 2010

Signed-off-by: Philipp Reisner <philipp.reisner@linbit.com>
Signed-off-by: Lars Ellenberg <lars.ellenberg@linbit.com>

48acf868

drbd: new configuration parameter c-min-rate · 0f0601f4

Lars Ellenberg authored Aug 11, 2010

We now track the data rate of locally submitted resync related requests,
and can thus detect non-resync activity on the lower level device.

If the current sync rate is above c-min-rate, and the lower level device
appears to be busy, we throttle the resyncer.
Signed-off-by: Philipp Reisner <philipp.reisner@linbit.com>
Signed-off-by: Lars Ellenberg <lars.ellenberg@linbit.com>

0f0601f4

drbd: reduce code duplication when receiving data requests · 80a40e43

Lars Ellenberg authored Aug 11, 2010

also canonicalize the return values of read_for_csum
and drbd_rs_begin_io to return -ESOMETHING, or 0 for success.
Signed-off-by: Philipp Reisner <philipp.reisner@linbit.com>
Signed-off-by: Lars Ellenberg <lars.ellenberg@linbit.com>

80a40e43

drbd: use rolling marks for resync speed calculation · 1d7734a0

Lars Ellenberg authored Aug 11, 2010

The current resync speed as displayed in /proc/drbd fluctuates a lot.
Using an array of rolling marks makes this calculation much more stable.
We used to have this (a long time ago with 0.7), but it got lost somehow.

If "stalled", do not discard the rest of the information, just add a
" (stalled)" tag to the progress line.

This patch also shortens a spinlock critical section somewhat, and
reduces the number of atomic operations in put_ldev.
Signed-off-by: Philipp Reisner <philipp.reisner@linbit.com>
Signed-off-by: Lars Ellenberg <lars.ellenberg@linbit.com>

1d7734a0

drbd: remove outdated comment and dead code · 0bb70bf6

Lars Ellenberg authored Aug 11, 2010

Signed-off-by: Philipp Reisner <philipp.reisner@linbit.com>
Signed-off-by: Lars Ellenberg <lars.ellenberg@linbit.com>

0bb70bf6

drbd: let drbd_free_ee implicitly free any digest · c36c3ced

Lars Ellenberg authored Aug 11, 2010

Signed-off-by: Philipp Reisner <philipp.reisner@linbit.com>
Signed-off-by: Lars Ellenberg <lars.ellenberg@linbit.com>

c36c3ced

drbd: Replaced some casts by an union. Improved comments · 85719573

Philipp Reisner authored Jul 21, 2010

Signed-off-by: Philipp Reisner <philipp.reisner@linbit.com>
Signed-off-by: Lars Ellenberg <lars.ellenberg@linbit.com>

85719573

drbd: Bugfix: rs_in_flight could become wrong if read_for_csum() requested reschedule later · d207450c
Philipp Reisner authored Jul 22, 2010
```
Signed-off-by: Philipp Reisner <philipp.reisner@linbit.com>
Signed-off-by: Lars Ellenberg <lars.ellenberg@linbit.com>
```
d207450c

drbd: The new, smarter resync speed controller · 778f271d

Philipp Reisner authored Jul 06, 2010

Signed-off-by: Philipp Reisner <philipp.reisner@linbit.com>
Signed-off-by: Lars Ellenberg <lars.ellenberg@linbit.com>

778f271d

drbd: New sync_param packet, that includes the parameters of the new controller · 8e26f9cc
Philipp Reisner authored Jul 06, 2010
```
Signed-off-by: Philipp Reisner <philipp.reisner@linbit.com>
Signed-off-by: Lars Ellenberg <lars.ellenberg@linbit.com>
```
8e26f9cc

drbd: New sync parameters for the smart resync rate controller · 9a31d716

Philipp Reisner authored Jul 05, 2010

Signed-off-by: Philipp Reisner <philipp.reisner@linbit.com>
Signed-off-by: Lars Ellenberg <lars.ellenberg@linbit.com>

9a31d716

drbd: fix list corruption (recent regression) · d28fd092

Lars Ellenberg authored Jul 09, 2010

The commit 288f422e
 drbd: Track all IO requests on the TL, not writes only
moved a list_add_tail(req, ) into a region where req
may have just been freed due to conflict detection.

Fix this by adding a proper cleanup section for that code path.
Signed-off-by: Philipp Reisner <philipp.reisner@linbit.com>
Signed-off-by: Lars Ellenberg <lars.ellenberg@linbit.com>

d28fd092

drbd: Initialize all members of sync_conf to their defaults [Bugz 315] · e756414f
Philipp Reisner authored Jun 29, 2010
```
Signed-off-by: Philipp Reisner <philipp.reisner@linbit.com>
Signed-off-by: Lars Ellenberg <lars.ellenberg@linbit.com>
```
e756414f
drbd: Make sure tl_restart(, resend) can not get called multiple times for a new connection · 67098930
Philipp Reisner authored Jun 24, 2010
```
Signed-off-by: Philipp Reisner <philipp.reisner@linbit.com>
Signed-off-by: Lars Ellenberg <lars.ellenberg@linbit.com>
```
67098930

drbd: Do not try to free tl_hash in drbd_disconnect() when IO is suspended · f70b3511

Philipp Reisner authored Jun 24, 2010

We may not free tl_hash when IO is suspended, since we can not wait
until ap_bio_cnt reaches zero.

We can do this after susp reched 0, since then tl_clear was called
Signed-off-by: Philipp Reisner <philipp.reisner@linbit.com>
Signed-off-by: Lars Ellenberg <lars.ellenberg@linbit.com>

f70b3511

drbd: Allow attach while IO is suspended · 8f488156

Philipp Reisner authored Jun 24, 2010

Signed-off-by: Philipp Reisner <philipp.reisner@linbit.com>
Signed-off-by: Lars Ellenberg <lars.ellenberg@linbit.com>

8f488156

drbd: Allow tl_restart() to do IO completion while IO is suspended · cfa03415

Philipp Reisner authored Jun 23, 2010

Signed-off-by: Philipp Reisner <philipp.reisner@linbit.com>
Signed-off-by: Lars Ellenberg <lars.ellenberg@linbit.com>

cfa03415

drbd: Fixed a deadlock, probably only affected UP machines · 84dfb9f5

Philipp Reisner authored Jun 23, 2010

After disconnect (most likely mdev->net_cnt == 0) and we are
still in an unstable state (!drbd_state_is_stable()). When we
get an IO request in drbd_get_max_buffers() (called from
__inc_ap_bio_cond(), called from inc_ap_bio()) we wake up
misc_wait. Misc_wait is also used in inc_ap_bio() to sleep
until the outcome of __inc_ap_bio_cond() changes. => Busy loop!

Solution: Have a dedicated wait queue for get_net_conf() and
put_net_conf().
Signed-off-by: Philipp Reisner <philipp.reisner@linbit.com>
Signed-off-by: Lars Ellenberg <lars.ellenberg@linbit.com>

84dfb9f5

drbd: Do not do a hard state change when establishing a connection [bugz 304] · 65d922c3

Philipp Reisner authored Jun 16, 2010

Make sure the state engine can deny two primaries to connect
Signed-off-by: Philipp Reisner <philipp.reisner@linbit.com>
Signed-off-by: Lars Ellenberg <lars.ellenberg@linbit.com>

65d922c3

drbd: Ensure that the peer was not rebootet in the meantime before resending TL · 481c6f50
Philipp Reisner authored Jun 22, 2010
```
Signed-off-by: Philipp Reisner <philipp.reisner@linbit.com>
Signed-off-by: Lars Ellenberg <lars.ellenberg@linbit.com>
```
481c6f50

drbd: Delayed creation of current-UUID · 43a5182c

Philipp Reisner authored Jun 11, 2010

When a fencing policy of "resource-and-stonith" is configured,
and DRBD looses connection to it's peer, we can delay the
creation of a new current-UUID until IO gets thawed.

That allows one to deploy fence-peer handlers that actually
commit suicide on the machine they get started.
Signed-off-by: Philipp Reisner <philipp.reisner@linbit.com>
Signed-off-by: Lars Ellenberg <lars.ellenberg@linbit.com>

43a5182c

drbd: Run the fence-peer helper asynchronously · 87f7be4c

Philipp Reisner authored Jun 11, 2010

Since we can not thaw the transfer log, the next logical step is
to allow reconnects while the fence-peer handler runs.
Signed-off-by: Philipp Reisner <philipp.reisner@linbit.com>
Signed-off-by: Lars Ellenberg <lars.ellenberg@linbit.com>

87f7be4c

drbd: Reduce the verbosity of some state transitions · 1616a254

Philipp Reisner authored Jun 10, 2010

State transitions in the space of non-allowed states used
to be very noisy. Reduce that, since that has little value
for the majority of the user base.
Signed-off-by: Philipp Reisner <philipp.reisner@linbit.com>
Signed-off-by: Lars Ellenberg <lars.ellenberg@linbit.com>

1616a254

drbd: Removing a by now obsolete clause in the state sanitizing · 999122bc

Philipp Reisner authored Jun 10, 2010

Signed-off-by: Philipp Reisner <philipp.reisner@linbit.com>
Signed-off-by: Lars Ellenberg <lars.ellenberg@linbit.com>

999122bc

drbd: Now we need to handle the ed_uuid of an diskless, unconnected primary correctly · 18a50fa2
Philipp Reisner authored Jun 21, 2010
```
Signed-off-by: Philipp Reisner <philipp.reisner@linbit.com>
Signed-off-by: Lars Ellenberg <lars.ellenberg@linbit.com>
```
18a50fa2
drbd: Disabled the crashed_primary detection for re-attach of last data while IO is frozen · 894c6a94
Philipp Reisner authored Jun 18, 2010
```
Signed-off-by: Philipp Reisner <philipp.reisner@linbit.com>
Signed-off-by: Lars Ellenberg <lars.ellenberg@linbit.com>
```
894c6a94
drbd: Do not allow a fencing-policy of resource-and-stonith with protocol A · 47ff2d0a
Philipp Reisner authored Jun 18, 2010
```
Signed-off-by: Philipp Reisner <philipp.reisner@linbit.com>
Signed-off-by: Lars Ellenberg <lars.ellenberg@linbit.com>
```
47ff2d0a

drbd: Finished the "on-no-data-accessible suspend-io;" functionality · 265be2d0

Philipp Reisner authored May 31, 2010

When no data is accessible (no connection to the peer, nor a local disk)
allow the user to select to freeze all IO operations instead of getting
IO errors.
Signed-off-by: Philipp Reisner <philipp.reisner@linbit.com>
Signed-off-by: Lars Ellenberg <lars.ellenberg@linbit.com>

265be2d0

drbd: Removed redundant error checks in the request code path · 905cd7d8

Philipp Reisner authored May 10, 2010

Signed-off-by: Philipp Reisner <philipp.reisner@linbit.com>
Signed-off-by: Lars Ellenberg <lars.ellenberg@linbit.com>

905cd7d8