Kernel - My Linux kernel repository

Age	Commit message (Collapse)	Author
2008-09-25	Btrfs: trivial sparse fixes	Christoph Hellwig
	Fix a bunch of trivial sparse complaints. Signed-off-by: Christoph Hellwig <hch@lst.de> Signed-off-by: Chris Mason <chris.mason@oracle.com>
2008-09-25	Btrfs: missing endianess conversion in insert_new_root	Christoph Hellwig
	Add two missing endianess conversions in this function, found by sparse. Signed-off-by: Christoph Hellwig <hch@lst.de> Signed-off-by: Chris Mason <chris.mason@oracle.com>
2008-09-25	Btrfs: Add a write ahead tree log to optimize synchronous operations	Chris Mason
	File syncs and directory syncs are optimized by copying their items into a special (copy-on-write) log tree. There is one log tree per subvolume and the btrfs super block points to a tree of log tree roots. After a crash, items are copied out of the log tree and back into the subvolume. See tree-log.c for all the details. Signed-off-by: Chris Mason <chris.mason@oracle.com>
2008-09-25	btrfs_search_slot: reduce lock contention by cowing in two stages	Chris Mason
	A btree block cow has two parts, the first is to allocate a destination block and the second is to copy the old bock over. The first part needs locks in the extent allocation tree, and may need to do IO. This changeset splits that into a separate function that can be called without any tree locks held. btrfs_search_slot is changed to drop its path and start over if it has to COW a contended block. This often means that many writers will pre-alloc a new destination for a the same contended block, but they cache their prealloc for later use on lower levels in the tree. Signed-off-by: Chris Mason <chris.mason@oracle.com>
2008-09-25	Btrfs: implement memory reclaim for leaf reference cache	Yan
	The memory reclaiming issue happens when snapshot exists. In that case, some cache entries may not be used during old snapshot dropping, so they will remain in the cache until umount. The patch adds a field to struct btrfs_leaf_ref to record create time. Besides, the patch makes all dead roots of a given snapshot linked together in order of create time. After a old snapshot was completely dropped, we check the dead root list and remove all cache entries created before the oldest dead root in the list. Signed-off-by: Chris Mason <chris.mason@oracle.com>
2008-09-25	Btrfs: Add a leaf reference cache	Yan Zheng
	Much of the IO done while dropping snapshots is done looking up leaves in the filesystem trees to see if they point to any extents and to drop the references on any extents found. This creates a cache so that IO isn't required. Signed-off-by: Chris Mason <chris.mason@oracle.com>
2008-09-25	Fix path slots selection in btrfs_search_forward	Yan
	We should decrease the found slot by one as btrfs_search_slot does when bin_search return 1 and node level > 0. Signed-off-by: Chris Mason <chris.mason@oracle.com>
2008-09-25	Btrfs: Create orphan inode records to prevent lost files after a crash	Josef Bacik
	Signed-off-by: Chris Mason <chris.mason@oracle.com>
2008-09-25	btrfs_next_leaf: do readahead when skip_locking is turned on	Chris Mason
	Signed-off-by: Chris Mason <chris.mason@oracle.com>
2008-09-25	Btrfs: Add locking around volume management (device add/remove/balance)	Chris Mason
	Signed-off-by: Chris Mason <chris.mason@oracle.com>
2008-09-25	Btrfs: Reduce contention on the root node	Chris Mason
	This calls unlock_up sooner in btrfs_search_slot in order to decrease the amount of work done with the higher level tree locks held. Also, it changes btrfs_tree_lock to spin for a big against the page lock before scheduling. This makes a big difference in context switch rate under highly contended workloads. Longer term, a better locking structure is needed than the page lock. Signed-off-by: Chris Mason <chris.mason@oracle.com>
2008-09-25	Btrfs: Online btree defragmentation fixes	Chris Mason
	The btree defragger wasn't making forward progress because the new key wasn't being saved by the btrfs_search_forward function. This also disables the automatic btree defrag, it wasn't scaling well to huge filesystems. The auto-defrag needs to be done differently. Signed-off-by: Chris Mason <chris.mason@oracle.com>
2008-09-25	Btrfs: Add btree locking to the tree defragmentation code	Chris Mason
	The online btree defragger is simplified and rewritten to use standard btree searches instead of a walk up / down mechanism. Signed-off-by: Chris Mason <chris.mason@oracle.com>
2008-09-25	Btrfs: Replace the transaction work queue with kthreads	Chris Mason
	This creates one kthread for commits and one kthread for deleting old snapshots. All the work queues are removed. Signed-off-by: Chris Mason <chris.mason@oracle.com>
2008-09-25	Btrfs: Fix snapshot deletion to release the alloc_mutex much more often.	Chris Mason
	This lowers the impact of snapshot deletion on the rest of the FS. Signed-off-by: Chris Mason <chris.mason@oracle.com>
2008-09-25	Btrfs: Add a skip_locking parameter to struct path, and make various funcs ↵	Chris Mason
	honor it Allocations may need to read in block groups from the extent allocation tree, which will require a tree search and take locks on the extent allocation tree. But, those locks might already be held in other places, leading to deadlocks. Since the alloc_mutex serializes everything right now, it is safe to skip the btree locking while caching block groups. A better fix will be to either create a recursive lock or find a way to back off existing locks while caching block groups. Signed-off-by: Chris Mason <chris.mason@oracle.com>
2008-09-25	Fix btrfs_next_leaf to check for new items after dropping locks	Chris Mason
	Signed-off-by: Chris Mason <chris.mason@oracle.com>
2008-09-25	Fix btrfs_del_ordered_inode to allow forcing the drop during unlinks	Chris Mason
	This allows us to delete an unlinked inode with dirty pages from the list instead of forcing commit to write these out before deleting the inode. Signed-off-by: Chris Mason <chris.mason@oracle.com>
2008-09-25	Drop locks in btrfs_search_slot when reading a tree block.	Chris Mason
	One lock per btree block can make for significant congestion if everyone has to wait for IO at the high levels of the btree. This drops locks held by a path when doing reads during a tree search. Signed-off-by: Chris Mason <chris.mason@oracle.com>
2008-09-25	Btrfs: Replace the big fs_mutex with a collection of other locks	Chris Mason
	Extent alloctions are still protected by a large alloc_mutex. Objectid allocations are covered by a objectid mutex Other btree operations are protected by a lock on individual btree nodes Signed-off-by: Chris Mason <chris.mason@oracle.com>
2008-09-25	Btrfs: Start btree concurrency work.	Chris Mason
	The allocation trees and the chunk trees are serialized via their own dedicated mutexes. This means allocation location is still not very fine grained. The main FS btree is protected by locks on each block in the btree. Locks are taken top / down, and as processing finishes on a given level of the tree, the lock is released after locking the lower level. The end result of a search is now a path where only the lowest level is locked. Releasing or freeing the path drops any locks held. Signed-off-by: Chris Mason <chris.mason@oracle.com>
2008-09-25	Btrfs: Allocator fix variety pack	Chris Mason
	* Force chunk allocation when find_free_extent has to do a full scan * Record the max key at the start of defrag so it doesn't run forever * Block groups might not be contiguous, make a forward search for the next block group in extent-tree.c * Get rid of extra checks for total fs size * Fix relocate_one_reference to avoid relocating the same file data block twice when referenced by an older transaction * Use the open device count when allocating chunks so that we don't try to allocate from devices that don't exist Signed-off-by: Chris Mason <chris.mason@oracle.com>
2008-09-25	Btrfs: Handle write errors on raid1 and raid10	Chris Mason
	When duplicate copies exist, writes are allowed to fail to one of those copies. This changeset includes a few changes that allow the FS to continue even when some IOs fail. It also adds verification of the parent generation number for btree blocks. This generation is stored in the pointer to a block, and it ensures that missed writes to are detected. Signed-off-by: Chris Mason <chris.mason@oracle.com>
2008-09-25	Btrfs: Pass down the expected generation number when reading tree blocks	Chris Mason
	Signed-off-by: Chris Mason <chris.mason@oracle.com>
2008-09-25	Btrfs: Fix balance_level to free the middle block if there is room in the ↵	Chris Mason
	left one balance level starts by trying to empty the middle block, and then pushes from the right to the middle. This might empty the right block and leave a small number of pointers in the middle. Signed-off-by: Chris Mason <chris.mason@oracle.com>
2008-09-25	Btrfs: Don't empty the middle buffer in push_nodes_for_insert	Chris Mason
	Signed-off-by: Chris Mason <chris.mason@oracle.com>
2008-09-25	Btrfs: Fix split_node to require more empty slots in the node as well	Chris Mason
	Signed-off-by: Chris Mason <chris.mason@oracle.com>
2008-09-25	Btrfs: Make sure nodes have enough room for a double split	Chris Mason
	Signed-off-by: Chris Mason <chris.mason@oracle.com>
2008-09-25	Btrfs: Don't wait on tree block writeback before freeing them anymore	Chris Mason
	This isn't required anymore because we don't reallocate blocks that have already been written in this transaction. Signed-off-by: Chris Mason <chris.mason@oracle.com>
2008-09-25	Btrfs: Add chunk uuids and update multi-device back references	Chris Mason
	Block headers now store the chunk tree uuid Chunk items records the device uuid for each stripes Device extent items record better back refs to the chunk tree Block groups record better back refs to the chunk tree The chunk tree format has also changed. The objectid of BTRFS_CHUNK_ITEM_KEY used to be the logical offset of the chunk. Now it is a chunk tree id, with the logical offset being stored in the offset field of the key. This allows a single chunk tree to record multiple logical address spaces, upping the number of bytes indexed by a chunk tree from 2^64 to 2^128. Signed-off-by: Chris Mason <chris.mason@oracle.com>
2008-09-25	Btrfs: Disable extra debugging checks on tree blocks	Chris Mason
	Signed-off-by: Chris Mason <chris.mason@oracle.com>
2008-09-25	Btrfs: Retry metadata reads in the face of checksum failures	Chris Mason
	Signed-off-by: Chris Mason <chris.mason@oracle.com>
2008-09-25	Btrfs: Do metadata checksums for reads via a workqueue	Chris Mason
	Before, metadata checksumming was done by the callers of read_tree_block, which would set EXTENT_CSUM bits in the extent tree to show that a given range of pages was already checksummed and didn't need to be verified again. But, those bits could go away via try_to_releasepage, and the end result was bogus checksum failures on pages that never left the cache. The new code validates checksums when the page is read. It is a little tricky because metadata blocks can span pages and a single read may end up going via multiple bios. Signed-off-by: Chris Mason <chris.mason@oracle.com>
2008-09-25	Change btrfs_map_block to return a structure with mappings for all stripes	Chris Mason
	Signed-off-by: Chris Mason <chris.mason@oracle.com>
2008-09-25	Btrfs: Properly dirty buffers in the split corner cases	Chris Mason
	Signed-off-by: Chris Mason <chris.mason@oracle.com>
2008-09-25	Btrfs: Verify checksums on tree blocks found without read_tree_block	Chris Mason
	Checksums were only verified by btrfs_read_tree_block, which meant the functions to probe the page cache for blocks were not validating checksums. Normally this is fine because the buffers will only be in cache if they have already been validated. But, there is a window while the buffer is being read from disk where it could be up to date in the cache but not yet verified. This patch makes sure all buffers go through checksum verification before they are used. This is safer, and it prevents modification of buffers before they go through the csum code. Signed-off-by: Chris Mason <chris.mason@oracle.com>
2008-09-25	Reorder the flags field in struct btrfs_header and record a flag on writeout	Chris Mason
	This allows detection of blocks that have already been written in the running transaction so they can be recowed instead of modified again. It is step one in trusting the transid field of the block pointers. Signed-off-by: Chris Mason <chris.mason@oracle.com>
2008-09-25	Btrfs: Add support for multiple devices per filesystem	Chris Mason
	Signed-off-by: Chris Mason <chris.mason@oracle.com>
2008-09-25	Call btrfs_cow_block while lowering tree level.	Yan
	When freeing root block of a tree, btrfs_free_extent' parameter 'ref_generation' is from root block itseft. When freeing non-root block, 'ref_generation' is from its parent. so when converting a non-root block to root block, we must guarantee its generation is equal to its parent's generation. Signed-off-by: Chris Mason <chris.mason@oracle.com>
2008-09-25	Btrfs: Copy correct tree when inserting into slot 0	Chris Mason
	Signed-off-by: Chris Mason <chris.mason@oracle.com>
2008-09-25	Btrfs: Add inode item and backref in one insert, reducing cpu usage	Chris Mason
	Signed-off-by: Chris Mason <chris.mason@oracle.com>
2008-09-25	Btrfs: During deletes and truncate, remove many items at once from the tree	Chris Mason
	Signed-off-by: Chris Mason <chris.mason@oracle.com>
2008-09-25	Btrfs: Add data=ordered support	Chris Mason
	This forces file data extents down the disk along with the metadata that references them. The current implementation is fairly simple, and just writes out all of the dirty pages in an inode before the commit. Signed-off-by: Chris Mason <chris.mason@oracle.com>
2008-09-25	Btrfs: Force inlining off in a few places to save stack usage	Chris Mason
	Signed-off-by: Chris Mason <chris.mason@oracle.com>
2008-09-25	Btrfs: Add readahead to the online shrinker, and a mount -o alloc_start= for ↵	Chris Mason
	testing Signed-off-by: Chris Mason <chris.mason@oracle.com>
2008-09-25	Btrfs: Less aggressive readahead on deletes	Chris Mason
	Signed-off-by: Chris Mason <chris.mason@oracle.com>
2008-09-25	kmalloc a few large stack objects in the btrfs_ioctl path	Chris Mason
	Signed-off-by: Chris Mason <chris.mason@oracle.com>
2008-09-25	Btrfs: Add mount option to turn off data cow	Chris Mason
	A number of workloads do not require copy on write data or checksumming. mount -o nodatasum to disable checksums and -o nodatacow to disable both copy on write and checksumming. In nodatacow mode, copy on write is still performed when a given extent is under snapshot. Signed-off-by: Chris Mason <chris.mason@oracle.com>
2008-09-25	Btrfs: Add back pointers from extents to the btree or file referencing them	Chris Mason
	Signed-off-by: Chris Mason <chris.mason@oracle.com>
2008-09-25	Btrfs: Implement generation numbers in block pointers	Chris Mason
	Signed-off-by: Chris Mason <chris.mason@oracle.com>