Commits · 335f7d4f22fd27ea86398a3617ce41ab3d478ae6 · Kirill Smelkov / linux

22 Oct, 2023 40 commits

bcachefs: clean up post-eof folios on -ENOSPC · 335f7d4f

Brian Foster authored Mar 29, 2023

The buffered write path batches folio creations in the file mapping
based on the requested size of the write. Under low free space
conditions, it is possible to add a bunch of folios to the mapping
and then return a short write or -ENOSPC due to lack of space. If
this occurs on an extending write, the file size is updated based on
the amount of data successfully written to the file. If folios were
added beyond the final i_size, they may hang around until reclaimed,
truncated or encountered unexpectedly by another operation.

For example, generic/083 reproduces a sequence of events where a
short write leaves around one or more post-EOF folios on an inode, a
subsequent zero range request extends beyond i_size and overlaps
with an aforementioned folio, and __bch2_truncate_folio() happens
across it and complains.

Update __bch2_buffered_write() to keep track of the start offset of
the last folio added to the mapping for a prospective write. After
i_size is updated, check whether this offset starts beyond EOF. If
so, truncate pagecache beyond the latest EOF to clean up any folios
that don't reside at least partially within EOF upon completion of
the write.
Signed-off-by: Brian Foster <bfoster@redhat.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

335f7d4f

bcachefs: fix truncate overflow if folio is beyond EOF · 4ad6aa46

Brian Foster authored Mar 29, 2023

generic/083 occasionally reproduces a panic caused by an overflow
when accessing the bch_folio_sector array of the folio being
processed by __bch2_truncate_folio(). The immediate cause of the
overflow is that the folio offset is beyond i_size, and therefore
the sector index calculation underflows on subtraction of the folio
offset.

One cause of this is mainly observed on nocow mounts. When nocow is
enabled, fallocate performs physical block allocation (as opposed to
block reservation in cow mode), which range_has_data() then
interprets as valid data that requires partial zeroing on truncate.
Therefore, if a post-eof zero range request lands across post-eof
preallocated blocks, __bch2_truncate_folio() may actually create a
post-eof folio in order to perform zeroing. To avoid this problem,
update range_has_data() to filter out unwritten blocks from folio
creation and partial zeroing.

Even though we should never create folios beyond EOF like this, the
mere existence of such folios is not necessarily a fatal error. Fix
up the truncate code to warn about this condition and not overflow
the sector array and possibly crash the system. The addition of this
warning without the corresponding unwritten extent fix has shown
that various other fstests are able to reproduce this problem fairly
frequently, but often in ways that doesn't necessarily result in a
kernel panic or a change in user observable behavior, and therefore
the problem goes undetected.
Signed-off-by: Brian Foster <bfoster@redhat.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

4ad6aa46

bcachefs: Enable large folios · 550a6a49
Kent Overstreet authored Mar 19, 2023
```
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>
```
550a6a49

bcachefs: Check for folios that don't have bch_folio attached · 34fdcf06

Kent Overstreet authored Mar 27, 2023

With large folios, it's now incidentally possible to end up with a
clean, uptodate folio in the page cache that doesn't have a bch_folio
attached, if a folio has to be split.

This patch fixes __bch2_truncate_folio() to check for this; other code
paths appear to handle it.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

34fdcf06

bcachefs: bch2_readahead() large folio conversion · 9567413c

Kent Overstreet authored Mar 17, 2023

Readahead now uses the new filemap_get_contig_folios_d() helper.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

9567413c

bcachefs: filemap_get_contig_folios_d() · 40022c01

Kent Overstreet authored Mar 23, 2023

Add a new helper for getting a range of contiguous folios and returning
them in a darray.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

40022c01

bcachefs: bch_folio_sector_state improvements · a1774a05

Kent Overstreet authored Mar 23, 2023

 - X-macro-ize the bch_folio_sector_state enum: this means we can easily
   generate strings, which is helpful for debugging.

 - Add helpers for state transitions: folio_sector_dirty(),
   folio_sector_undirty(), folio_sector_reserve()

 - Add folio_sector_set(), a single helper for changing folio sector
   state just so that we have a single place to instrument when we're
   debugging.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

a1774a05

bcachefs: bch2_truncate_page() large folio conversion · 959f7368
Kent Overstreet authored Mar 19, 2023
```
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>
```
959f7368
bcachefs: bch2_buffered_write large folio conversion · c42b57c4
Kent Overstreet authored Mar 18, 2023
```
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>
```
c42b57c4
bcachefs: bch_folio can now handle multi-order folios · 49fe78ff
Kent Overstreet authored Mar 17, 2023
```
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>
```
49fe78ff

bcachefs: More assorted large folio conversion · 33e2eb96

Kent Overstreet authored Mar 17, 2023

Various misc small conversions in fs-io.c for large folios.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

33e2eb96

bcachefs: bch2_seek_pagecache_data() folio conversion · a86a92cb

Kent Overstreet authored Mar 19, 2023

This converts bch2_seek_pagecache_data() to handle large folios.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

a86a92cb

bcachefs: bch2_seek_pagecache_hole() folio conversion · e8d28c3e

Kent Overstreet authored Mar 19, 2023

This converts bch2_seek_pagecache_hole() to handle large folios.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

e8d28c3e

bcachefs: bio_for_each_segment_all() -> bio_for_each_folio_all() · ff9c301f
Kent Overstreet authored Mar 19, 2023
```
This converts the writepage end_io path to folios.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>
```
ff9c301f

bcachefs: Initial folio conversion · 30bff594

Kent Overstreet authored Mar 17, 2023

This converts fs-io.c to pass folios, not pages. We're not handling
large folios yet, there's no functional changes in this patch - just a
lot of churn doing the initial type conversions.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

30bff594

bcachefs: Rename bch_page_state -> bch_folio · 3342ac13
Kent Overstreet authored Mar 17, 2023
```
Start of the large folio conversion.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>
```
3342ac13

bcachefs: Add a bch_page_state assert · c437e153

Kent Overstreet authored Mar 27, 2023

Seeing an odd bug with page/folio state not being properly initialized,
this is to help track it down.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

c437e153

bcachefs: Add a cond_resched() call to journal_keys_sort() · 27763692

Kent Overstreet authored Apr 04, 2023

We're just doing cpu work here and it could take awhile, a
cond_resched() is definitely needed.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

27763692

bcachefs: Improve trace_move_extent_fail() · bb6c4b92

Kent Overstreet authored Mar 10, 2023

This greatly expands the move_extent_fail tracepoint - now it includes
all the information we have available, including exactly why the extent
wasn't updated.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

bb6c4b92

bcachefs: Print out counters correctly · 86691994

Kent Overstreet authored Mar 30, 2023

Most counters aren't in units of sectors, and the ones that are should
just be switched to bytes, for simplicity.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

86691994

bcachefs: Add missing bch2_err_class() call · dde72e18

Kent Overstreet authored Mar 30, 2023

We're not supposed to return our private error codes to userspace.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

dde72e18

bcachefs: Rip out code for storing backpointers in alloc keys · 62a03559

Kent Overstreet authored Mar 31, 2023

We don't store backpointers in alloc keys anymore, since we gained the
btree write buffer.

This patch drops support for backpointers in alloc keys, and revs the on
disk format version so that we know a fsck is required.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

62a03559

bcachefs: use reservation for log messages during recovery · 349b1d83

Brian Foster authored Mar 22, 2023

If we block on journal reservation attempting to log journal
messages during recovery, particularly for the first message(s)
before we start doing actual work, chances are the filesystem ends
up deadlocked.

Allow logged messages to use reserved journal space to mitigate this
problem. In the worst case where no space is available whatsoever,
this at least allows the fs to recognize that the journal is stuck
and fail the mount gracefully.
Signed-off-by: Brian Foster <bfoster@redhat.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

349b1d83

bcachefs: Improve trans_restart_split_race tracepoint · 3d86f13d

Kent Overstreet authored Mar 30, 2023

Seeing occasional test failures where we get stuck in a livelock that
involves this event - this will help track it down.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

3d86f13d

bcachefs: Data update path no longer leaves cached replicas · 25d8f405

Kent Overstreet authored Mar 29, 2023

It turns out that it's currently impossible to invalidate buckets
containing only cached data if they're part of a stripe. The normal
bucket invalidate path can't do it because we have to be able to
incerement the bucket's gen, which isn't correct becasue it's still a
member of the stripe - and the bucket invalidate path makes the bucket
availabel for reuse right away, which also isn't correct for buckets in
stripes.

What would work is invalidating cached data by following backpointers,
except that cached replicas don't currently get backpointers - because
they would be awkward for the existing bucket invalidate path to delete
and they haven't been needed elsewhere.

So for the time being, to prevent running out of space in stripes,
switch the data update path to not leave cached replicas; we may revisit
this in the future.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

25d8f405

bcachefs: Rhashtable based buckets_in_flight for copygc · 32de2ea0

Kent Overstreet authored Mar 11, 2023

Previously, copygc used a fifo for tracking buckets in flight - this had
the disadvantage of being fixed size, since we pass references to
elements into the move code.

This restructures it to be a hash table and linked list, since with
erasure coding we need to be able to pipeline across an arbitrary number
of buckets.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

32de2ea0

bcachefs: Use BTREE_ITER_INTENT in ec_stripe_update_extent() · 6bdefe9c

Kent Overstreet authored Mar 29, 2023

This adds a flags param to bch2_backpointer_get_key() so that we can
pass BTREE_ITER_INTENT, since ec_stripe_update_extent() is updating the
extent immediately.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

6bdefe9c

bcachefs: move snapshot_t to subvolume_types.h · 4f77dcde
Kent Overstreet authored Mar 29, 2023
```
this doesn't need to be in bcachefs.h
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>
```
4f77dcde

bcachefs: Fix bch2_get_key_or_hole() · 1546cf97

Kent Overstreet authored Mar 28, 2023

This fixes an off by one error, due to confusing closed vs. half open
intervals.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

1546cf97

bcachefs: Check return code from need_whiteout_for_snapshot() · 2a6c302f

Kent Overstreet authored Mar 28, 2023

This could return a transaction restart; we need to check for that.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

2a6c302f

bcachefs: bch2_dev_freespace_init() Print out status every 10 seconds · e9b9e475

Kent Overstreet authored Mar 22, 2023

It appears freespace init can still take awhile, and we've had a report
or two of it getting stuck - let's have it print out where it's at every
10 seconds.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

e9b9e475

bcachefs: Run freespace init in device hot add path · b1c945b3

Kent Overstreet authored Mar 22, 2023

Like in the recovery, and device add, we have to check if devices don't
have the freespace btree initialized - this was missed in the device hot
add path.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

b1c945b3

bcachefs: Improved copygc wait debugging · 0fb11e08

Kent Overstreet authored Mar 17, 2023

This just adds a line for how long copygc has been waiting to sysfs
copygc_wait, helpful for debugging why copygc isn't running.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

0fb11e08

bcachefs: Call bch2_path_put_nokeep() before bch2_path_put() · 11f11737

Kent Overstreet authored Mar 21, 2023

bch2_path_put_nokeep() is sketchy, and we should consider removing it:
it unconditionally frees btree_paths once their ref hits 0.

The assumption is that we only use it for paths that have never been
visible outside the btree core btree code; i.e. higher level code will
never be making assumptions about locking based on these paths.

However, there's subtle brokenness with this approach:

 - If we call bch2_path_put(), then bch2_path_put_nokeep(),
   bch2_path_put() may free the first path on the assumption that we we
   have another path keeping a node locked - but then
   bch2_path_put_nokeep() just unconditionally frees it.

The same bug may arise if we're calling bch2_path_put() and
bch2_path_put_nokeep() on the same (refcounted) path, or two adjacent
paths that point to the same btree node.

This patch hacks around one of these bugs by calling
bch2_path_put_nokeep() first in bch2_trans_iter_exit.
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

11f11737

bcachefs: drop unnecessary journal stuck check from space calculation · 030e9f92

Brian Foster authored Mar 21, 2023

The journal stucking check in bch2_journal_space_available() is
particularly aggressive and can lead to premature shutdown in some
rare cases. This is difficult to reproduce, but also comes along
with a fatal error and so is worthwhile to be cautious.

For example, we've seen instances where the journal is under heavy
reservation pressure, the journal allocation path transitions into
the final available journal bucket, the journal write path
immediately consumes that bucket and calls into
bch2_journal_space_available(), which then in turn flags the journal
as stuck because there is no available space and shuts down the
filesystem instead of submitting the journal write (that would have
otherwise succeeded).

To avoid this problem, simplify the journal stuck checking by just
relying on the higher level logic in the journal reservation path.
This produces more useful debug output and is a more reliable
indicator that things have bogged down.
Signed-off-by: Brian Foster <bfoster@redhat.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

030e9f92

bcachefs: refactor journal stuck checking into standalone helper · db1bf729

Brian Foster authored Mar 21, 2023

bcachefs checks for journal stuck conditions both in the journal
space calculation code and the journal reservation slow path. The
logic in both places is rather tricky and can result in
non-deterministic failure characteristics and debug output.

In preparation to condense journal stuck handling to a single place,
refactor the __journal_res_get() logic into a standalone helper.
Since multiple callers into the reservation code can result in
duplicate reports, use the ->err_seq field as a serialization
mechanism for the debug dump. Finally, add some comments to help
explain the logic and hopefully facilitate further improvements in
the future.
Signed-off-by: Brian Foster <bfoster@redhat.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

db1bf729

bcachefs: gracefully unwind journal res slowpath on shutdown · 23fd4f4d

Brian Foster authored Mar 20, 2023

bcachefs detects journal stuck conditions in a couple different
places. If the logic in the journal reservation slow path happens to
detect the problem, I've seen instances where the filesystem remains
deadlocked even though it has been shut down. This is occasionally
reproduced by generic/333, and usually manifests as one or more
tasks stuck in the journal reservation slow path.

To help avoid this problem, repeat the journal error check in
__journal_res_get() once under spinlock to cover the case where the
previous lock holder might have triggered shutdown. This also helps
avoid spurious/duplicate stuck reports. Also, wake the journal from
the halt code to make sure blocked callers of the journal res
slowpath have a chance to wake up and observe the pending error.
This survives an overnight looping run of generic/333 without the
aforementioned lockups.
Signed-off-by: Brian Foster <bfoster@redhat.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

23fd4f4d

bcachefs: more aggressive fast path write buffer key flushing · 873555f0

Brian Foster authored Mar 17, 2023

The btree write buffer flush code is prone to causing journal
deadlock due to inefficient use and release of reservation space.
Reservation is not pre-reserved for write buffered keys (as is done
for key cache keys, for example), because the write buffer flush
side uses a fast path that attempts insertion without need for any
reservation at all.

The write buffer flush attempts to deal with this by inserting keys
using the BTREE_INSERT_JOURNAL_RECLAIM flag to return an error on
journal reservations that require blocking. Upon first error, it
falls back to a slow path that inserts in journal order and supports
moving the associated journal pin forward.

The problem is that under pathological conditions (i.e. smaller log,
larger write buffer and journal reservation pressure), we've seen
instances where the fast path fails fairly quickly without having
completed many insertions, and then the slow path is unable to push
the journal pin forward enough to free up the space it needs to
completely flush the buffer. This problem is occasionally reproduced
by fstest generic/333.

To avoid this problem, update the fast path algorithm to skip key
inserts that fail due to inability to acquire needed journal
reservation without immediately breaking out of the loop. Instead,
insert as many keys as possible, zap the sequence numbers to mark
them as processed, and then fall back to the slow path to process
the remaining set in journal order. This reduces the amount of
journal reservation that might be required to flush the entire
buffer and increases the odds that the slow path is able to move the
journal pin forward and free up space as keys are processed.
Signed-off-by: Brian Foster <bfoster@redhat.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

873555f0

bcachefs: use dedicated workqueue for tasks holding write refs · 8bff9875

Brian Foster authored Mar 23, 2023

A workqueue resource deadlock has been observed when running fsck
on a filesystem with a full/stuck journal. fsck is not currently
able to repair the fs due to fairly rapid emergency shutdown, but
rather than exit gracefully the fsck process hangs during the
shutdown sequence. Fortunately this is easily recoverable from
userspace, but the root cause involves code shared between the
kernel and userspace and so should be addressed.

The deadlock scenario involves the main task in the bch2_fs_stop()
-> bch2_fs_read_only() path waiting on write references to drain
with the fs state lock held. A bch2_read_only_work() workqueue task
is scheduled on the system_long_wq, blocked on the state lock.
Finally, various other write ref holding workqueue tasks are
scheduled to run on the same workqueue and must complete in order to
release references that the initial task is waiting on.

To avoid this problem, we can split the dependent workqueue tasks
across different workqueues. It's a bit of a waste to create a
dedicated wq for the read-only worker, but there are several tasks
throughout the fs that follow the pattern of acquiring a write
reference and then scheduling to the system wq. Use a local wq
for such tasks to break the subtle dependency between these and the
read-only worker.
Signed-off-by: Brian Foster <bfoster@redhat.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

8bff9875

bcachefs: remove unused bch2_trans_log_msg() · 76c70c57

Brian Foster authored Mar 22, 2023

Signed-off-by: Brian Foster <bfoster@redhat.com>
Signed-off-by: Kent Overstreet <kent.overstreet@linux.dev>

76c70c57