Commits · bb125baf512bffef19c510f1c53353a378537070 · Kirill Smelkov / linux

22 Oct, 2023 40 commits

bcachefs: Delete warning from promote_alloc() · bb125baf

Kent Overstreet authored Jun 04, 2023

It's possible to see a -BCH_ERR_ENOSPC_disk_reservation here, and that's
fine.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

bb125baf

bcachefs: Fix bch2_fsck_ask_yn() · 4f2c166e

Kent Overstreet authored Jun 04, 2023

 - getline() output includes a newline, without stripping that we were
   just looping

 - Make the prompt clearer
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

4f2c166e

bcachefs: replicas_deltas_realloc() uses allocate_dropping_locks() · 21da6101
Kent Overstreet authored May 28, 2023
```
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>
```
21da6101

bcachefs: Convert acl.c to allocate_dropping_locks() · 5ff10c0a

Kent Overstreet authored May 28, 2023

More work to avoid allocating memory with btree locks held.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

5ff10c0a

bcachefs: allocate_dropping_locks() · d95dd378

Kent Overstreet authored May 28, 2023

Add two new helpers for allocating memory with btree locks held: The
idea is to first try the allocation with GFP_NOWAIT|__GFP_NOWARN, then
if that fails - unlock, retry with GFP_KERNEL, and then call
trans_relock().
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

d95dd378

bcachefs: Use unlikely() in bch2_err_matches() · 3ebfc8fe
Kent Overstreet authored May 29, 2023
```
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>
```
3ebfc8fe

bcachefs: Fix error handling in promote path · 4c4a8f20

Kent Overstreet authored May 29, 2023

The promote path had a BUG_ON() for unknown error type, which we're now
seeing: change it to a WARN_ON() - because we're curious what this is -
and otherwise handle it in the normal error path.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

4c4a8f20

bcachefs: fs-io: Eliminate GFP_NOFS usage · 5718fda0

Kent Overstreet authored May 28, 2023

GFP_NOFS doesn't ever make sense. If we're allocatingc memory it should
be GFP_NOWAIT if btree locks are held, GFP_KERNEL otherwise.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

5718fda0

bcachefs: bch2_trans_kmalloc no longer allocates memory with btree locks held · 78367aaa

Kent Overstreet authored May 28, 2023

When allocating memory, gfp flags should generally be

 - GFP_NOWAIT|__GFP_NOWARN if btree locks are held
 - GFP_NOFS if in the IO path or otherwise holding resources needed for
   IO submission
 - GFP_KERNEL otherwise
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

78367aaa

bcachefs: drop_locks_do() · b5fd7566

Kent Overstreet authored May 28, 2023

Add a new helper for the common pattern of:
 - trans_unlock()
 - do something
 - trans_relock()
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

b5fd7566

bcachefs: GFP_NOIO -> GFP_NOFS · 19c304be

Kent Overstreet authored May 28, 2023

GFP_NOIO dates from the bcache days, when we operated under the block
layer. Now, GFP_NOFS is more appropriate, so switch all GFP_NOIO uses to
GFP_NOFS.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

19c304be

bcachefs: Ensure bch2_btree_node_get() calls relock() after unlock() · e1d29c5f

Kent Overstreet authored May 28, 2023

Fix a bug where bch2_btree_node_get() might call bch2_trans_unlock() (in
fill) without calling bch2_trans_relock(); this is a bug when it's done
in the core btree code.

Also, twea bch2_btree_node_mem_alloc() to drop btree locks before doing
a blocking memory allocation.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

e1d29c5f

bcachefs: Avoid __GFP_NOFAIL · 70d41c9e

Kent Overstreet authored May 28, 2023

We've been using __GFP_NOFAIL for allocating struct bch_folio, our
private per-folio state.

However, that struct is variable size - it holds state for each sector
in the folio, and folios can be quite large now, which means it's
possible for bch_folio to be larger than PAGE_SIZE now.

__GFP_NOFAIL allocations are undesirable in normal circumstances, but
particularly so at >= PAGE_SIZE, and warnings are emitted for that.

So, this patch adds proper error paths and eliminates most uses of
__GFP_NOFAIL. Also, do some more cleanup of gfp flags w.r.t. btree node
locks: we can use GFP_KERNEL, but only if we're not holding btree locks,
and if we are holding btree locks we should be using GFP_NOWAIT.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

70d41c9e

bcachefs: Fix corruption with writeable snapshots · ad520141

Kent Overstreet authored May 27, 2023

When partially overwriting an extent in an older snapshot, the existing
extent has to be split.

If the existing extent was overwritten in a different (sibling)
snapshot, we have to ensure that the split won't be visible in the
sibling snapshot.

data_update.c already has code for this,
bch2_insert_snapshot_writeouts() - we just need to move it into
btree_update_leaf.c and change bch2_trans_update_extent() to use it as
well.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

ad520141

bcachefs: Convert -ENOENT to private error codes · e47a390a

Kent Overstreet authored May 27, 2023

As with previous conversions, replace -ENOENT uses with more informative
private error codes.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

e47a390a

bcachefs: trans_for_each_path_safe() · f154c3eb

Kent Overstreet authored May 27, 2023

bch2_btree_trans_to_text() is used on btree_trans objects that are owned
by different threads - when printing out deadlock cycles - so we need a
safe version of trans_for_each_path(), else we race with seeing a
btree_path that was just allocated and not fully initialized:
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

f154c3eb

bcachefs: Fix a quota read bug · e7ffda56

Kent Overstreet authored May 27, 2023

bch2_fs_quota_read() could see an inode that's been deleted
(KEY_TYPE_inode_generation) - bch2_fs_quota_read_inode() needs to check
for that instead of erroring.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

e7ffda56

bcachefs: Fix move_extent_fail counter · c26463ce

Kent Overstreet authored May 26, 2023

fail counters need to be events, not numbers of sectors - or the
calculations the tests use for determining if we've had too many
slowpath events don't work.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

c26463ce

bcachefs: Don't reuse reflink btree keyspace · fc0ee376

Kent Overstreet authored May 25, 2023

We've been seeing difficult to debug "missing indirect extent" bugs,
that fsck doesn't seem to find.

One possibility is that there was a missing indirect extent, but then a
new indirect extent was created at the location of the previous indirect
extent.

This patch eliminates that possibility by always creating new indirect
extents right after the last one, at the end of the reflink btree.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

fc0ee376

mean and variance: Add a missing include · db32bb9a
Kent Overstreet authored Jun 04, 2023
```
abs() is in math.h
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>
```
db32bb9a

mean and variance: More tests · 65bc4109

Kent Overstreet authored May 25, 2023

Add some more tests that test conventional and weighted mean
simultaneously, and with a table of values that represents events that
we'll be using this to look for so we can verify-by-eyeball that the
output looks sane.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

65bc4109

six locks: Disable percpu read lock mode in userspace · aab5e097

Kent Overstreet authored Jun 10, 2023

When running in userspace, we currently don't have a real percpu
implementation available - at least in bcachefs-tools, which is where
this code is currently used in userspace.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

aab5e097

six locks: Use atomic_try_cmpxchg_acquire() · 2d9200cf

Kent Overstreet authored May 25, 2023

This switches to a newer cmpxchg variant which updates @old for us on
failure, simplifying the cmpxchg loops a bit and supposedly generating
better code.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

2d9200cf

six locks: Fix an unitialized var · c4687a4a

Kent Overstreet authored May 25, 2023

In the conversion to atomic_t, six_lock_slowpath() ended up calling
six_lock_wakeup() in the failure path with a state variable that was
never initialized - whoops.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

c4687a4a

six locks: Delete redundant comment · 96e53e90
Kent Overstreet authored May 23, 2023
```
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>
```
96e53e90
six locks: Tiny bit more tidying · 2ab62310
Kent Overstreet authored May 22, 2023
```
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>
```
2ab62310
six locks: Seq now only incremented on unlock · 32913f49
Kent Overstreet authored Jun 16, 2023
```
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>
```
32913f49
six locks: Split out seq, use atomic_t instead of atomic64_t · 2804d0f1
Kent Overstreet authored May 22, 2023
```
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>
```
2804d0f1

six locks: Single instance of six_lock_vals · a4e9e1f0

Kent Overstreet authored Jun 16, 2023

Since we're not generating different versions of the lock functions for
each lock type, the constant propagation we were trying to do before is
no longer useful - this is now a small code size decrease.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

a4e9e1f0

six_locks: Kill test_bit()/set_bit() usage · 357c1261

Kent Overstreet authored May 22, 2023

This deletes the crazy cast-atomic-to-unsigned-long, and replaces them
with atomic_and() and atomic_or().
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

357c1261

six locks: lock->state.seq no longer used for write lock held · b60c8e9e

Kent Overstreet authored Jun 16, 2023

lock->state.seq is shortly being moved out of lock->state, to kill the
depedency on atomic64; in preparation for that, we change the write
locking bit to write locked.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

b60c8e9e

six locks: Simplify six_relock() · dc88b65f

Kent Overstreet authored Jun 16, 2023

The next patch is going to move lock->seq out of lock->state. This
replaces six_relock() with a much simpler implementation based on
trylock.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

dc88b65f

six locks: Improve spurious wakeup handling in pcpu reader mode · 37f612be
Kent Overstreet authored May 21, 2023
```
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>
```
37f612be

six locks: Documentation, renaming · 91d16f16

Kent Overstreet authored May 21, 2023

 - Expanded and revamped overview documentation in six.h, giving an
   overview of all features
 - docbook-comments for all external interfaces
 - Rename some functions for simplicity, i.e.
   six_lock_ip_type() -> six_lock_ip()
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

91d16f16

six locks: Kill six_lock_state union · 1fb4fe63

Kent Overstreet authored May 20, 2023

As suggested by Linus, this drops the six_lock_state union in favor of
raw bitmasks.

On the one hand, bitfields give more type-level structure to the code.
However, a significant amount of the code was working with
six_lock_state as a u64/atomic64_t, and the conversions from the
bitfields to the u64 were deemed a bit too out-there.

More significantly, because bitfield order is poorly defined (#ifdef
__LITTLE_ENDIAN_BITFIELD can be used, but is gross), incrementing the
sequence number would overflow into the rest of the bitfield if the
compiler didn't put the sequence number at the high end of the word.

The new code is a bit saner when we're on an architecture without real
atomic64_t support - all accesses to lock->state now go through
atomic64_*() operations.

On architectures with real atomic64_t support, we additionally use
atomic bit ops for setting/clearing individual bits.

Text size: 7467 bytes -> 4649 bytes - compilers still suck at
bitfields.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

1fb4fe63

six locks: Simplify dispatch · c4bd3491

Kent Overstreet authored May 20, 2023

Originally, we used inlining/flattening to cause the compiler to
generate different versions of lock/trylock/relock/unlock for each lock
type - read, intent, and write. This made the individual functions
smaller and let the compiler eliminate table lookups: however, as the
code has gotten more complicated these optimizations have gotten less
worthwhile, and all the tricky inlining and dispatching made the code
less readable.

Text size: 11015 bytes -> 7467 bytes, and benchmarks show no loss of
performance.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

c4bd3491

six locks: Centralize setting of waiting bit · d2c86b77

Kent Overstreet authored May 20, 2023

Originally, the waiting bit was always set by trylock() on failure:
however, it's now set by __six_lock_type_slowpath(), with wait_lock held
- which is the more correct place to do it.

That made setting the waiting bit in trylock redundant, so this patch
deletes that.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

d2c86b77

six locks: Remove hacks for percpu mode lost wakeup · 0157f9c5

Kent Overstreet authored May 21, 2023

The lost wakeup bug hasn't been observed in awhile, and we're trying to
provoke it and determine if it still exists.

This patch removes some defenses that were added to attempt to track it
down; if it still exists, this should make it easier to see it.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

0157f9c5

six locks: Kill six_lock_pcpu_(alloc|free) · 0d2234a7

Kent Overstreet authored May 20, 2023

six_lock_pcpu_alloc() is an unsafe interface: it's not safe to allocate
or free the percpu reader count on an existing lock that's in use, the
only safe time to allocate percpu readers is when the lock is first
being initialized.

This patch adds a flags parameter to six_lock_init(), and instead of
six_lock_pcpu_free() we now expose six_lock_exit(), which does the same
thing but is less likely to be misused.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

0d2234a7

six locks: six_lock_readers_add() · 01bf56a9

Kent Overstreet authored May 20, 2023

This moves a helper out of the bcachefs code that shouldn't have been
there, since it touches six lock internals.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

01bf56a9