Commits · 4bd4035e64c2a90b9c939135e95be0106205b370 · Kirill Smelkov / linux

22 Oct, 2023 40 commits

bcachefs: Handle sb buffer resizing in __copy_super() · 4bd4035e

Kent Overstreet authored Feb 12, 2023

This fixes a rare buffer overrun when one field is growing and another
field is shrinking - and is a nice simplification as well.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

4bd4035e

bcachefs: Fix failure to read btree roots · 806c8a6a

Kent Overstreet authored Feb 12, 2023

If failed to read a btree root - or if we're not using a btree root,
because of the reconstruct_alloc option - make sure we update the
corresponding info for the key/level for the root on disk.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

806c8a6a

bcachefs: Don't run triggers when repairing in __bch2_mark_reflink_p() · 32770815

Daniel Hill authored Feb 12, 2023

Triggers current trip-up on the faulty reflink we're trying to repair,
Disabling them lets us fix broken reflink and continue.
Signed-off-by: Daniel Hill <daniel@gluo.nz>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

32770815

bcachefs: let __bch2_btree_insert() pass in flags · 8ffa11a2

Daniel Hill authored Jan 20, 2023

This patch is prep work for the following patch.
Signed-off-by: Daniel Hill <daniel@gluo.nz>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

8ffa11a2

bcachefs: Improve locking in __bch2_set_nr_journal_buckets() · 76966dbf

Kent Overstreet authored Feb 11, 2023

This refactors to not call bch2_journal_block() with c->sb_lock held.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

76966dbf

bcachefs: More info on check_bucket_ref() error · c1f59ef6
Kent Overstreet authored Feb 11, 2023
```
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>
```
c1f59ef6
bcachefs: Add missing include · 930c0c4c
Kent Overstreet authored Feb 11, 2023
```
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>
```
930c0c4c
bcachefs: Handle btree node rewrites before going RW · a1f26d70
Kent Overstreet authored Feb 11, 2023
```
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>
```
a1f26d70
bcachefs: Nocow locking fixup · 09d70d0b
Kent Overstreet authored Feb 11, 2023
```
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>
```
09d70d0b
bcachefs: Add some logging for btree node rewrites due to errors · 12795a19
Kent Overstreet authored Feb 10, 2023
```
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>
```
12795a19

bcachefs: Ensure btree node cache is not more than half dirty · 637de729

Kent Overstreet authored Nov 11, 2021

Tweak journal reclaim to ensure the btree node cache isn't more
than half dirty so that memory reclaim can always make progress - the
same as we do for the btree key cache.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>

637de729

bcachefs: Add max nr of IOs in flight to the move path · c782c583
Kent Overstreet authored Jan 09, 2023
```
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>
```
c782c583
bcachefs: Add an assert to bch2_bucket_nocow_unlock() · 01efebd8
Kent Overstreet authored Jan 26, 2023
```
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>
```
01efebd8

bcachefs: don't block reads if we're promoting · 3482dd6a

Daniel Hill authored Jan 06, 2023

The promote path calls data_update_init() and now that we take locks here,
there's potential for promote to block our read path, just error
when we can't take the lock instead of blocking.
Signed-off-by: Daniel Hill <daniel@gluo.nz>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

3482dd6a

bcachefs: Fix promote path leak · 0093b9e9
Kent Overstreet authored Jan 05, 2023
```
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>
```
0093b9e9

bcachefs: Improve invalidate_one_bucket() error messages · 629a21b6

Kent Overstreet authored Jan 03, 2023

Make sure to check for lru entries that point to buckets that don't
exist as well as buckets in the wrong state, and improve the error
message we print out.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

629a21b6

bcachefs: Fix move_ctxt_wait_event() · 46eea9cb

Kent Overstreet authored Jan 03, 2023

We shouldn't be evaluating cond again if it already returned true.

This fixes a bug when this helper is used for taking nocow locks.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

46eea9cb

bcachefs: Fix deadlock on nocow locks in data move path · 7ffb6a7e

Kent Overstreet authored Jan 02, 2023

The recent nocow locking rework introduced a deadlock in the data move
path: the new nocow locking scheme uses a hash table with a fixed size
array for chaining, meaning on hash collision we may have to wait for
other locks to be released before we can lock a bucket.

And since the data move path needs to submit writes from the same thread
that's taking nocow locks and submitting reads, this introduces a
deadlock.

This shouldn't happen often in practice, but since the data move path
can keep large numbers of IOs in flight simultaneously, it's something
we have to handle.

This patch makes move_ctxt_wait_event() available to
bch2_data_update_init() and uses it when appropriate, which is our
normal solution to this kind of thing.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

7ffb6a7e

bcachefs: BKEY_INVALID_FROM_JOURNAL · dbe17f18
Kent Overstreet authored Dec 20, 2022
```
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>
```
dbe17f18
bcachefs: Change bkey_invalid() rw param to flags · facafdcb
Kent Overstreet authored Dec 20, 2022
```
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>
```
facafdcb

bcachefs: Rework lru btree · 83f33d68

Kent Overstreet authored Dec 05, 2022

This patch changes how the LRU index works:

Instead of using KEY_TYPE_lru where the bucket the lru entry points to
is part of the value, this switches to KEY_TYPE_set and encoding the
bucket we refer to in the low bits of the key.

This means that we no longer have to check for collisions when inserting
LRU entries. We'll be making using of this in the next patch, which adds
a btree write buffer - a pure write buffer for btree updates, where
updates are appended to a simple array and then periodically sorted and
batch inserted.

This is a new on disk format version, and a forced upgrade.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

83f33d68

bcachefs: Improved nocow locking · 350175bf

Kent Overstreet authored Dec 14, 2022

This improves the nocow lock table so that hash table entries have
multiple locks, and locks specify which bucket they're for - i.e. we can
now resolve hash collisions.

This is important because the allocator has to skip buckets that are
locked in the nocow lock table, and previously hash collisions would
cause it to spuriously skip unlocked buckets.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

350175bf

bcachefs: handle failed data_update_init cleanup · f3a37e76

Daniel Hill authored Dec 09, 2022

data_update_init allocates several resources, but we forget to clean
these up when it fails.
Signed-off-by: Daniel Hill <daniel@gluo.nz>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

f3a37e76

bcachefs: expose nocow_lock table in sysfs · 71fe1465

Daniel Hill authored Dec 07, 2022

Signed-off-by: Daniel Hill <daniel@gluo.nz>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

71fe1465

bcachefs: bucket_gens btree · 5250b74d

Kent Overstreet authored Nov 25, 2022

To improve mount times, add a btree for just bucket gens, 256 of them
per key: this means we'll have to scan drastically less metadata at
startup.

This adds
 - trigger for keeping it in sync with the all btree
 - initialization code, for filesystems from previous versions
 - new path for reading bucket gens
 - new fsck code

And a new on disk format version.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

5250b74d

bcachefs: Inline bch2_two_state_(trylock|unlock) · 19fe87e0

Kent Overstreet authored Nov 23, 2022

Standard inlining of fast paths - these locks are now used by our new
nocow mode.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

19fe87e0

bcachefs: Nocow support · a8b3a677

Kent Overstreet authored Nov 02, 2022

This adds support for nocow mode, where we do writes in-place when
possible. Patch components:

 - New boolean filesystem and inode option, nocow: note that when nocow
   is enabled, data checksumming and compression are implicitly disabled

 - To prevent in-place writes from racing with data moves
   (data_update.c) or bucket reuse (i.e. a bucket being reused and
   re-allocated while a nocow write is in flight, we have a new locking
   mechanism.

   Buckets can be locked for either data update or data move, using a
   fixed size hash table of two_state_shared locks. We don't have any
   chaining, meaning updates and moves to different buckets that hash to
   the same lock will wait unnecessarily - we'll want to watch for this
   becoming an issue.

 - The allocator path also needs to check for in-place writes in flight
   to a given bucket before giving it out: thus we add another counter
   to bucket_alloc_state so we can track this.

 - Fsync now may need to issue cache flushes to block devices instead of
   flushing the journal. We add a device bitmask to bch_inode_info,
   ei_devs_need_flush, which tracks devices that need to have flushes
   issued - note that this will lead to unnecessary flushes when other
   codepaths have already issued flushes, we may want to replace this with
   a sequence number.

 - New nocow write path: look up extents, and if they're writable write
   to them - otherwise fall back to the normal COW write path.

XXX: switch to sequence numbers instead of bitmask for devs needing
journal flush

XXX: ei_quota_lock being a mutex means bch2_nocow_write_done() needs to
run in process context - see if we can improve this
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

a8b3a677

bcachefs: Data update support for unwritten extents · 4dcd1cae

Kent Overstreet authored Nov 14, 2022

The data update path requires special support for unwritten extents - we
still need to be able to move them, but there's no need to read or write
anything.

This patch adds a new error code to tell bch2_move_extent() that we're
short circuiting the read, and adds bch2_update_unwritten_extent() to
create a reservation then call __bch2_data_update_index_update().
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

4dcd1cae

bcachefs: Unwritten extents support · 79203111

Kent Overstreet authored Nov 13, 2022

 - bch2_extent_merge checks unwritten bit
 - read path returns 0s for unwritten extents without actually reading
 - reflink path skips over unwritten extents
 - bch2_bkey_ptrs_invalid() checks for extents with both written and
   unwritten extents, and non-normal extents (stripes, btree ptrs) with
   unwritten ptrs
 - fiemap checks for unwritten extents and returns
   FIEMAP_EXTENT_UNWRITTEN
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

79203111

bcachefs: bch2_extent_update_i_size_sectors() · 2f1f7fe9

Kent Overstreet authored Nov 14, 2022

In the io path, when we do the extent update we also have to update the
inode - for i_size and i_sectors updates, as well as for bi_journal_seq
for fsync.

This factors that out into a new helper which will be used in the new
nocow mode, in the unwritten extent conversion path.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

2f1f7fe9

bcachefs: bch2_extent_fallocate() · 70de7a47

Kent Overstreet authored Nov 13, 2022

This factors out part of __bchfs_fallocate() in fs-io.c into an new,
lower level io.c helper, which creates a single extent reservation.

This is prep work for nocow support - the new helper will shortly gain
the ability to create unwritten extents.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

70de7a47

bcachefs: Skip inode unpack/pack in bch2_extent_update() · 9bcbc030

Kent Overstreet authored Oct 21, 2022

This takes advantage of the new inode type to skip the expensive
pack/unpack when inode updates are required in the extent update path.
Additionally, we now skip the inode update entirely when i_sectors and
i_size aren't changing.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

9bcbc030

bcachefs: Drop old maybe_extending optimization · b08b492e

Kent Overstreet authored Nov 08, 2021

The extend update path had an optimization to avoid updating the inode
if we knew we were definitely not extending the file. But now that we're
updating inodes on every extent update - for fsync - that code can be
deleted.
Signed-off-by: Kent Overstreet <kent.overstreet@gmail.com>

b08b492e

bcachefs: KEY_TYPE_inode_v3, metadata_version_inode_v3 · 8dd69d9f

Kent Overstreet authored Oct 21, 2022

Move bi_size and bi_sectors into the non-varint portion of the inode, so
that the write path can update them without going through the relatively
expensive unpack/pack operations.

Other changes:
 - Add a field for the offset of the varint section, so we can add new
   non-varint fields without needing a new inode type, like alloc_v3
 - Move bi_mode into the flags field, so that the varint section can be
   u64 aligned
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

8dd69d9f

bcachefs: Start snapshots before bch2_gc() · 47b323a0

Kent Overstreet authored Jan 19, 2023

bch2_gc may require snapshots to be started - the repair path when
checking the reflink btree may do updates to the extents btree.

This moves bch2_fs_initialize_subvolumes() and bch2_fs_snapshots_start()
to before bch2_gc() - since we haven't gone RW yet, the updates in
bch2_fs_initialize_subvolumes() are done via the journal replay keys
list, so it's fine to do this before bch2_gc().
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

47b323a0

bcachefs: Improve bch2_check_alloc_info() · d23124c7

Kent Overstreet authored Nov 30, 2022

This factors out a new helper from bch2_dev_freespace_init(),
bch2_get_key_or_hole(), and uses it in bch2_check_alloc_info(): we're
now able to process holes in the alloc btree as ranges, instead of one
bucket at a time.

This will improve fsck performance on new filesystems, or filesystems
where not every bucket has been used yet.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

d23124c7

bcachefs: Improve bch2_dev_freespace_init() · cc65f565

Kent Overstreet authored Nov 26, 2022

This makes bch2_dev_freespace_init() much faster: instead of processing
every bucket on the device one at a time, we handle ranges of missing
keys all at once: the freespace btree is an extents style btree, so we
only have to insert one freespace key for every range of missing keys
in the alloc btree.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

cc65f565

fixup bcachefs: New on disk format: Backpointers · 7c057d35
Kent Overstreet authored Feb 12, 2023
```
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>
```
7c057d35

bcachefs: Don't use key cache during fsck · 53b1c6f4

Kent Overstreet authored Oct 14, 2022

The btree key cache mainly helps with lock contention, at the cost of
additional memory overhead. During some fsck passes the memory overhead
really matters, but fsck is single threaded so lock contention is an
issue - so skipping the key cache during fsck will help with
performance.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

53b1c6f4

bcachefs: Run check_extents_to_backpointers() in multiple passes · b32f9a57

Kent Overstreet authored Sep 28, 2022

Similer to the previous patch for check_backpointers_to_extents(), if
the alloc + backpointers btrees do not fit in ram we need to run into
multiple passes.

The counting of btree nodes that fit in memory is different here,
because we have to walk the alloc and backpointers btrees at the same
time, since a backpointer could reside in either of them and we don't
know which without checking both.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

b32f9a57