DNE3 Performance Improvements
Author: Di Wang
Ticket: LU-20808 "Reduce DNE Cross-MDT RPCs for DNE operation"
Date: 2026-09-23
Summary
Lustre Distributed Namespace Environment (DNE) distributes directories and metadata across multiple Metadata Targets (MDTs) to scale aggregate metadata throughput. However, cross-MDT operations require acquiring remote locks, fetching remote metadata, and executing distributed transactions.
This overhead significantly impacts metadata-heavy workloads, such as AI reinforcement learning sandboxes, code-agent evaluations, and multi-node compilation environments. These workloads frequently create small files, materialize Git worktrees, unpack virtual environments, and perform atomic renames. Redundant MDT-to-MDT round trips per request compound metadata load and tail latency across the cluster.
This proposal focuses on high-rate, small-file lifecycles and metadata preparation phases. Bulk I/O workloads like streaming reads and checkpointing remain outside its scope.
When processing a modifying operation in DNE, the coordinator MDT executes four distinct phases:
- Discover and validate objects: Fetch remote attributes, xattrs (LinkEA, ACLs, layouts), and look up directory entries for permission checks, lock ordering, and validation.
- Acquire distributed locks: Enqueue remote LDLM inode-bit and PDO locks. Granting UPDATE or XATTR locks invalidates the local OSP cache.
- Execute distributed transaction: Commit local OSD updates and pack remote updates into a single compound
OUT_UPDATERPC per participating MDT, logging update records for recovery. - Retire recovery records asynchronously: The commit thread cancels update-log records in the background after subtransactions commit. This background cleanup does not contribute to foreground client latency.
While Phase 3 batches updates and Phase 4 runs asynchronously, Phase 1 currently issues multiple synchronous OUT_UPDATE RPCs for attributes, xattrs, and lookups.
This serialized fetch phase creates a major latency bottleneck.
The table below summarizes foreground and background OUT_UPDATE RPCs for representative operations on an unmodified tree during the reint handler window:
| Operation | Cross-MDT Sync Reads | Foreground Tx RPCs | Delayed Log Cleanup | OUT_UPDATE Total |
LDLM Enqueues | Handler Duration |
|---|---|---|---|---|---|---|
mkdir with a remote parent |
8 | 1 | 1 | 10 | 2 | 9.5 ms |
rmdir with a remote parent |
2 | 1 | 1 | 4 | 3 | 6.1 ms |
| cross-MDT directory rename | 3 | 5 | 1 | 9 | 1 | 20.2 ms |
| cross-MDT file rename | 9 | 1 | 1 | 11 | 3 | 8.5 ms |
Lock enqueue counts depend on which MDT coordinates the operation: renames coordinated on MDT0000 take the Big Filesystem Lock (BFL) locally, whereas renames coordinated on other MDTs add a fourth remote enqueue RPC.
Categorizing preparation reads shows recurring, redundant fetch patterns across operations.
mkdir with a remote parent (the parent here is the filesystem root):
| # | Operation | Caller |
|---|---|---|
| 1 | attr_get |
mdt_check_resource_ids(), before the parent lock
|
| 2 | lookup |
lockless pre-lock lookup (LU-10235) |
| 3 | attr_get |
mdd_lookup() permission check, after the lock
|
| 4 | lookup |
authoritative lookup under the lock |
| 5 | lookup |
mdd_create_sanity_check()
|
| 6 | xattr_get |
mdd_acl_init() — default ACL
|
| 7 | xattr_get |
mdd_pin_init() — pin
|
| 8 | xattr_get |
lod_get_default_lov_striping() via lod_ah_init() — default LOV
|
Note that the target name is looked up three separate times, and reads 3 through 8 occur sequentially after acquiring parent locks.
rmdir with remote parent: Executes two reads after lock acquisition: an attr_get for mdd_lookup() permission validation and the lookup itself.
Both sit immediately after the parent's name-hash enqueue and both belong to a single mdo_lookup(), which is the same shape the create and rename parent intents already handle.
A separate trace of a file unlink shows an identical pattern, so one intent covers both.
See Implementation sequence item 10.
Cross-MDT file rename: Involves synchronous preparation reads across both the source parent directory and the source object:
| # | Object | Operation | Caller |
|---|---|---|---|
| 1 | source parent | xattr_get |
mdd_parent_fid() via mdd_is_subdir() — linkEA
|
| 2 | source parent | xattr_get |
mdt_attr_get_pfid() — the same linkEA, read again
|
| 3 | source parent | xattr_get |
mdt_stripe_get() — LMV
|
| 4 | source parent | attr_get |
mdd_lookup() permission check, after the lock
|
| 5 | source parent | xattr_get |
lod_striping_load() — LMV
|
| 6 | source parent | lookup |
lod_index_try()
|
| 7 | source | attr_get |
lod_object_init() instantiation
|
| 8 | source | attr_get |
mdd_rename()
|
| 9 | source | xattr_get |
mdd_linkea_prepare() — linkEA
|
Read 2 duplicates Read 1 because mdd_parent_fid() invalidates the object cache; Reads 4–6 are serialized inside a single mdo_lookup(); and Read 8 duplicates Read 7 because acquiring the source lock flushes the OSP cache.
Cross-MDT directory rename: Requires only three preparation reads (attributes, LinkEA, and LMV), but issues five foreground transaction RPCs due to update-log object allocation, making it the slowest operation (20.2 ms).
MDTest results below illustrate DNE overhead, 10 iterations each.
Setting max-inherit-rr=2 keeps most directory operations local to a single MDT.
Setting max-inherit-rr=10 forces most directory operations across MDTs.
| Operation | Mean rate, single MDT (ops/s) | Mean rate, cross-MDT (ops/s) | Mean latency, RR=2 (ms/op) | Mean latency, RR=10 (ms/op) | Latency ratio, RR=10 / RR=2 |
|---|---|---|---|---|---|
| Directory creation | 30,154 | 4,066 | 0.269 | 2.000 | 7.4x slower |
| Directory rename | 36,876 | 5,151 | 0.217 | 1.589 | 7.3x slower |
| Directory removal | 21,913 | 3,444 | 0.366 | 2.390 | 6.5x slower |
| Directory stat | 86,893 | 16,491 | 0.093 | 0.488 | 5.2x slower |
| File stat | 37,868 | 29,708 | 0.212 | 0.270 | 1.3x |
| Tree removal | 78 | 75 | 0.013 | 0.014 | 1.1x |
| File removal | 21,038 | 20,277 | 0.381 | 0.395 | 1.0x |
| File read | 20,233 | 19,748 | 0.396 | 0.406 | 1.0x |
| File rename | 36,267 | 36,143 | 0.222 | 0.222 | 1.0x |
| File creation | 19,023 | 19,488 | 0.428 | 0.411 | 1.0x |
| File cross rename | 506 | 535 | 31.654 | 29.953 | 0.9x |
- Every directory operation costs 5x to 7x more latency once it crosses MDTs. Creation and rename are worst at 7.4x and 7.3x, removal 6.5x, and even stat - which performs no modification - 5.2x.
- File operations are essentially unaffected, all within 1.0x to 1.3x. Spreading files over two MDTs neither helps nor hurts them measurably, which is what makes the directory figures above attributable to cross-MDT preparation rather than to the benchmark or the hardware.
- Tree creation is excluded from the comparison: its run-to-run variance exceeds 100% (see the appendix), so no ratio drawn from it would be meaningful.
One result does not fit the pattern and is worth stating separately. File cross rename costs about 30 ms/op in both configurations - 140x more than an ordinary file rename in the same run - and crossing MDTs makes no difference to it. Its cost is therefore not cross-MDT preparation at all but the filesystem-wide rename lock, which a cross-directory file rename takes whether or not the two directories live on the same MDT. It is addressed separately, under Measured RPC reduction and performance.
In summary, cross-MDT layouts scale file operations well, but directory operations incur substantial inter-MDT RPC overhead.
To eliminate this overhead, we introduce semantic MDT-to-MDT lock intents for operation preparation.
A cross-MDT operation must enqueue a lock on the remote target regardless; the intent piggybacks the operation's synchronous preparation reads onto that enqueue.
The target MDT performs the lookup, permission checks, validation, and inheritance against its own local storage while granting the lock, returning the results in the lock reply.
Reads that would otherwise require separate OUT_UPDATE RPCs are eliminated, removing serialized round trips.
Existing locking, transaction, replay, and recovery semantics remain unchanged.
The same mechanism covers create and mkdir, rename, as well as unlink and rmdir.
A few reads fall outside it—some occur before a lock exists to carry an intent, and others must be reissued after lock acquisition—and are addressed separately.
The proposal below outlines the design, and the measured impact is presented under Measured RPC reduction and performance.
Proposal
Design Overview
For each modifying request, the client selects one MDT as the coordinator based on the operation and target object FIDs. The request flows through the MDT, MDD, and LOD layers. LOD routes operations on local objects to the local OSD, while operations on remote objects go through an OSP proxy. Foreground MDT-to-MDT work consists of three main phases:
- Remote lock acquisition: Sending LDLM enqueue RPCs for objects managed by remote MDTs.
- Remote preparation reads: Issuing synchronous
OUT_UPDATERPCs to fetch attributes, xattrs, and directory lookups. These operations are strictly serialized because each result determines whether and how subsequent requests are issued. - Distributed transaction execution: Committing local changes via the local OSD and packing remote updates into a single compound foreground
OUT_UPDATERPC per participating MDT.
After subtransactions commit, recovery-record cancellation may send an additional asynchronous OUT_UPDATE RPC.
Although this step consumes network and server CPU resources, it runs in the background outside the client critical path.
Operations involving a remote parent directory require two sequential locks: a whole-directory lock (typically LCK_CW) to prevent splitting and restriping, followed by a name-hash lock (typically LCK_PW) to serialize access to the specific child entry.
Rather than merging these lock enqueues, this proposal piggybacks operation preparation onto the second lock enqueue RPC.
Coordinator MDT Parent MDT
--------------- ----------
enqueue whole-directory CW lock ----------> grant lock
<------------ reply
enqueue name-hash PW lock
+ DNE create-prepare intent ----------> grant lock
run parent-side checks
derive child state
<------------ lock + prepare result
The lock intent is operation-aware rather than a generic prefetch request:
IT_DNE_CREATE_PREPARE = 0x00010000, IT_DNE_RENAME_PREPARE = 0x00020000, IT_DNE_UNLINK_PREPARE = 0x00040000,
Instead of sending raw parent attributes and xattrs back to the coordinator, the parent MDT performs the necessary checks locally while granting the lock.
For creation operations, the parent MDT verifies parent directory validity, performs the authoritative child lookup, checks permissions (MAY_WRITE | MAY_EXEC) using the caller's credentials, validates layout versions and creation constraints, applies SGID and project inheritance, processes default ACLs, and evaluates pin and striping policies (LMV/LOV).
Intermediate checks stay local to the parent MDT—only the overall status and essential initialization data for the child are returned.
Rename operations use the same intent structure with a smaller payload.
When locking the source parent, the intent passes the target entry name, and the reply returns the resolved FID alongside the parent attributes and LMV layout—replacing the three synchronous reads inside mdo_lookup().
When locking the source object itself, no name is passed (MCPR_NO_LOOKUP), and the reply returns only the object's attributes required by mdd_rename().
Because rename requires a subset of create data, it shares the same request/reply wire structures.
Unlink and rmdir operations use an even simpler payload. The intent attached to the parent lock enqueue passes the entry name and receives the resolved FID and parent attributes. The coordinator then evaluates deletion permissions locally using those returned attributes. Like rename, unlink reuses the common request/reply structures without defining dedicated types.
Certain parent metadata parameters must be returned to initialize the child object created on the coordinator MDT:
| Result | Why the coordinator needs it |
|---|---|
| Effective child mode and GID | SGID and ACL processing may change them |
| Effective project ID and flags | The child may inherit PROJINHERIT
|
| Existing child FID | Required for EEXIST, replay, or partial-create handling
|
| Access ACL | Written onto the new child when a default ACL exists |
| Default ACL | Inherited when the new child is a directory |
| Pin value | Inherited by a new file or directory |
| Effective LMV/LOV policy | Used to initialize the child's layout or placement |
For standard files and directories without default ACLs, pins, or custom striping policies, the returned payload remains lightweight.
How the MDT changes
To prevent logic duplication and code drift between local and remote execution paths, parent-side validation logic is refactored from the MDD create path into a single helper function. This helper accepts the parent object, child name, and proposed child attributes, returning the prepared result. For local operations, the coordinator invokes this helper directly after acquiring local locks. For remote operations, the remote MDT executes the same helper during lock processing and returns the result in the lock reply. Namespace insertion, parent timestamp updates, child object creation, and transaction execution remain unchanged.
For each operation the change is confined to the point where the parent lock is taken.
Create and mkdir: The lock acquisition step includes the preparation intent request. For local parents, the lock is acquired and the helper executes inline. For remote parents, the intent is attached to the name-hash lock enqueue. Upon success, the caller receives the granted lock alongside the resolved FID and child attributes, skipping redundant local lookups and attribute derivations.
Rename: Two intents are attached to separate lock enqueues: one on the source parent's name-hash lock (returning the source FID, parent attributes, and striping) and one on the source object lock (returning object attributes). The first intent result directly satisfies the lookup step, while the second populates the OSP object cache transparently.
Unlink and rmdir: A single intent attached to the parent name-hash lock returns the resolved FID and parent attributes. The coordinator checks removal permissions against these cached attributes without additional remote RPCs.
Across all operations, lock granting is decoupled from intent outcome: the requested lock is granted regardless of whether preparation succeeds.
This separation is critical to prevent lock leaks. Current locking APIs return non-zero errors only when lock acquisition fails, triggering cleanup routines that assume no locks are held. If a lock function returned an error due to a preparation failure while holding the granted lock, that lock would leak.
If preparation fails or produces incomplete data, the lock response reports success with an empty result payload. The coordinator then falls back smoothly to standard lock-protected lookups and RPC reads. Furthermore, because the coordinator maintains full client group credentials while intent requests carry limited supplementary groups, this fallback ensures security checks remain exact.
For creation operations, lockless pre-lock lookups (used to detect existing entries without taking write locks) are bypassed when targeting remote parents. Because remote lockless lookups require synchronous RPCs and preparation performs an authoritative lookup under lock anyway, skipping the pre-lock check saves one synchronous round trip for all successful creations.
OSP changes
The OSP remote-lock layer normally issues standard LDLM enqueue RPCs. When a prepare descriptor is attached and the remote target advertises support, OSP updates its handling to:
- Allocate the new intent request format.
- Pack the intent opcode, target name, proposed child attributes, caller credentials, and reply buffer bounds.
- Set
LDLM_FL_HAS_INTENT. - Call the existing synchronous
ldlm_cli_enqueue(). - Unpack the per-operation status and prepared child state before releasing the request.
For create and unlink operations, the intent is attached exclusively to the second (name-hash) UPDATE lock enqueue, leaving the initial whole-directory lock unchanged. For rename operations, intents attach to both the source parent name-hash lock and the source object lock.
Returned state and the OSP object cache
Intent replies return two result types: lookup outcomes (consumed directly by the coordinator) and object metadata (parent attributes, ACLs, pins, and layouts), which is populated directly into OSP's object cache.
Populating the OSP cache keeps subsystem changes minimal. Higher-level callers (such as MDD reading ACLs or LOD reading striping) access data through standard read routines without requiring intent-aware interfaces.
Caching intent-returned metadata maintains strict consistency. Acquiring an UPDATE or XATTR lock normally invalidates the local OSP cache to drop stale pre-lock data. However, metadata returned in a lock intent reply is fetched by the remote target under the exact lock being granted, guaranteeing freshness equivalent to a post-lock read.
To maintain correctness, standard post-lock cache invalidation is suppressed when an intent reply arrives. OSP invalidates stale cache entries immediately after these entries are accessed.
Internal lock interface
The optional prepare descriptor passes from MDT through MDD and LOD down to OSP.
It uses a dedicated intent field in struct ldlm_enqueue_info to avoid conflicting with ei_cbdata (used for callbacks and striped directory state):
/* optional variable-length results an intent reply may carry */
enum md_lock_intent_buf {
MLI_BUF_PATTR = 0,
MLI_BUF_DEF_ACL = 1,
MLI_BUF_PIN = 2,
MLI_BUF_LMV = 3,
MLI_BUF_LOV = 4,
MLI_BUF_DEF_LMV = 5,
MLI_BUF_MAX
};
struct md_lock_intent {
__u64 mli_opc; /* enum ldlm_intent_flags */
const void *mli_request; /* intent-specific request body */
size_t mli_request_size;
void *mli_reply; /* intent-specific reply body */
size_t mli_reply_size;
const char *mli_name; /* entry name the intent operates on */
size_t mli_namelen;
/* destinations for the variable-length results, provided by the
* caller: lb_len is the inline limit on the way out and the length
* actually returned on the way back
*/
struct lu_buf mli_reply_buf[MLI_BUF_MAX];
unsigned int mli_replied:1; /* set when the reply was filled in */
};
struct ldlm_enqueue_info {
...
struct md_lock_intent *ei_intent;
};
Callers without intents set ei_intent to NULL.
MDD and LOD pass the enqueue info transparently, while OSP serializes the intent context for remote requests.
Remote MDT changes
The remote MDT intent dispatcher routes incoming requests to dedicated handlers sharing a unified wire format. Handlers grant the requested lock, reconstruct user credentials, execute the shared validation helper against local storage, and return status along with derived child attributes. Replayed and resent lock enqueues are processed consistently with existing intent handlers.
Locks are granted regardless of preparation status. If validation fails or optional attributes overflow response buffers, the reply flags the omission and the coordinator falls back to standard reads while retaining lock ownership.
Intent handlers are strictly read-only and create no transactions or durable updates. All namespace mutations and recovery logging remain managed by the subsequent distributed transaction.
Two key operational rules govern handler execution:
Lock parameters: Handlers grant the exact lock resources and policy modes requested by the coordinator rather than recalculating them locally, preventing mismatches caused by differing local node configurations.
Credential handling: Handlers adopt user identities directly from incoming RPC contexts without re-evaluating identity mappings, avoiding authorization discrepancies between MDTs. Handlers accept intent requests only over secure MDT-to-MDT connections.
Rename intent handlers support two modes: entry mode (resolves target entry names, checks lookup permissions, and returns FIDs, attributes, and striping) and object mode (operates directly on target objects and returns object attributes).
Wire protocol and interoperability
Wire protocol additions consist of one intent opcode per operation type alongside unified request and reply payloads. A bitmask flags which optional fields are present in each message.
Replies contain fixed-size fields for parent attributes and dynamic fields for returned xattrs. Output size is strictly bounded by requester-specified buffer limits.
Rename and unlink operations reuse create wire structures, setting only relevant fields and flags.
Intent support is negotiated as an MDT-to-MDT connection feature flag. If a peer lacks intent support (e.g., during rolling upgrades), the coordinator seamlessly falls back to standard lock enqueues and separate RPC reads.
Credentials and locking correctness
Because authorization checks execute on the parent MDT, intent requests must transmit full caller context—including effective UID/GID, group memberships, capabilities, and umask. Remote target nodes execute checks strictly under the caller's identity.
Prepared results remain protected by locks held across the preparation and transaction execution phases, covering child entry existence, parent attributes, permissions, ACLs, pins, and layout policies.
Locking ordering follows standard DNE rules to prevent deadlocks. All local reads performed by target intent handlers access local storage directly, avoiding recursive OSP RPCs.
Replay, errors, and fallback
Intent handlers are read-only. On request resends or replays, the target MDT re-evaluates the intent under existing lock protection.
Intent replies indicate three possible processing outcomes:
- Incomplete preparation: Reported as success with an empty result payload. The coordinator proceeds with standard lookups and reads under acquired locks.
- Advisory results: Processed similarly using fallback lookup paths.
- Missing or truncated xattrs: Indicated by clear bitmask flags. The coordinator uses returned attributes and fetches only the missing item individually.
Transaction replays bypass intents entirely. Replayed creations require full lockless lookup and version checks, acquiring locks without intents. Replayed renames and unlinks execute explicit name resolution to preserve version bookkeeping.
If intents are unsupported, unpackable, or incomplete, the system defaults unconditionally to standard DNE workflows without altering operation semantics.
Measured RPC reduction and performance
A prototype has been built to verify the idea, and the figures below are taken from it rather than estimated.
It covers mkdir and rename only for now.
Everything reported here should therefore be read as evidence that the approach works and is worth completing, not as the result of the finished design.
RPC counts
Measured from single-operation debug traces, one before-and-after pair per operation:
| Operation | OUT_UPDATE before |
after | Sync reads before | after |
|---|---|---|---|---|
mkdir with a remote parent |
7 | 3 | 5 | 1 |
| cross-MDT file rename, target does not exist | 12 | 7 | 10 | 5 |
| cross-MDT file rename, target exists | 16 | 9 | 14 | 7 |
The LDLM enqueue count is unchanged in every case - two for mkdir, four for rename - because the intents ride on enqueues that were already being issued.
The mkdir pair was taken on a different filesystem state from the survey in the Summary, which is why its absolute counts differ from the 8 reads recorded there; each pair is internally consistent, before against after on the same tree.
For mkdir the post-lock parent reads go to zero: the lookup, the permission check, the default ACL, the pin and the default layout all arrive in the lock reply.
For the target-exists rename, every removed read is accounted for as well: the duplicate linkEA read; the parent's attributes, striping and lookup, all three of which belong to a single lookup call; the source's attributes; and two parent attribute reads that the old cache invalidation had stranded.
Confirmation in the traces is direct - there is no lookup RPC anywhere in the patched log, and each removed read appears instead as a cache hit.
Not every read can be merged into an enqueue, which is why the counts do not fall to the transaction alone. Some happen before any lock exists to attach an intent to: an object's attributes are fetched when it is first instantiated, well before the operation reaches its locking phase. Others have to be reissued after a lock is granted, precisely because the grant is what made the earlier answer unusable - the re-check taken under the rename lock is the clearest case, since obtaining a fresh answer is the entire point of it. Those remain ordinary synchronous reads.
IOPS
mdtest with two MDTs.
The baseline is the unmodified tree; the prototype carries the mkdir and rename intents.
Two directory layouts were measured, because they stress the design differently: all tasks working in one shared directory, and one directory per task.
5 iterations each.
| Operation | baseline (ops/s) | prototype (ops/s) | change | |
|---|---|---|---|---|
| Directory creation | 1,015 ± 50 | 2,058 ± 92 | +103% | significant |
| File cross rename | 296 ± 4 | 399 ± 30 | +35% | significant |
| Directory rename | 1,740 ± 53 | 1,969 ± 70 | +13% | significant |
| Directory removal | 1,143 ± 33 | 1,211 ± 42 | +6% | within noise |
| File read | 16,056 ± 837 | 16,295 ± 1,027 | +1% | within noise |
| File rename | 5,420 ± 177 | 5,456 ± 287 | +1% | within noise |
| File creation | 4,826 ± 241 | 4,831 ± 385 | 0% | within noise |
| File stat | 41,081 ± 1,724 | 38,484 ± 3,262 | -6% | within noise |
| File removal | 3,147 ± 123 | 2,847 ± 533 | -10% | within noise |
| Directory stat | 20,050 ± 1,397 | 15,475 ± 846 | -23% | significant |
Tree creation and removal are omitted; their variance makes any comparison drawn from them meaningless.
The controls make this run readable. File creation, file read and file rename are all within 1.5%, and none of them is touched by an intent, so there is no drift between the two runs to discount. Whatever else moved, moved because of the change.
Directory creation doubles, which is the result the design predicts most directly: a shared parent is remote for half the tasks, and every such create drops from 7 RPCs to 3.
File cross rename gains 35%, and directory rename 13%. Cross rename remains slow in absolute terms because its cost is dominated by the filesystem-wide rename lock rather than by preparation reads, which the next section takes up.
Directory removal is unchanged, and it is the one operation here whose intent has not been built. That is the control the rest of the table should be judged against.
Directory stat loses 23%, and this is a genuine regression. Though directory stats seems varies a lot, since the stats mostly are memory operation in this test, and I will verify this further in a more scale test environment.
One directory per task
10 iterations each.
| Operation | baseline (ops/s) | prototype (ops/s) | change |
|---|---|---|---|
| Directory creation | 4,066 ± 566 | 5,657 ± 1,534 | +39% |
| File cross rename | 535 ± 20 | 664 ± 25 | +24% |
| File rename | 36,143 ± 1,777 | 41,859 ± 2,155 | +16% |
| File removal | 20,277 ± 580 | 23,242 ± 857 | +15% |
| Directory rename | 5,151 ± 848 | 5,809 ± 1,543 | +13% |
| File creation | 19,488 ± 566 | 21,923 ± 833 | +12% |
| File read | 19,748 ± 798 | 22,063 ± 1,421 | +12% |
| Directory stat | 16,491 ± 1,423 | 16,721 ± 1,891 | +1% |
| Directory removal | 3,444 ± 663 | 3,465 ± 692 | +1% |
Appendix
Master with current master
Single MDT operation (with --max-inherit-rr=2):
SUMMARY rate (in ops/sec): (of 10 iterations) Operation Max Min Mean Std Dev --------- --- --- ---- ------- Directory creation 36545.220 25034.586 30154.402 3896.568 Directory stat 98784.804 77697.116 86893.120 7385.451 Directory rename 38471.338 35097.273 36876.413 1189.058 Directory removal 22711.213 19945.653 21912.575 871.136 File creation 20463.339 12756.831 19023.351 2255.510 File stat 40990.573 33739.459 37868.077 2211.085 File read 21496.298 18611.273 20233.337 811.341 File rename 39783.351 29319.802 36266.547 3020.317 File cross rename 528.074 491.083 505.670 10.929 File removal 22486.627 19972.187 21038.295 880.323 Tree creation 31.514 1.707 8.087 8.699 Tree removal 101.491 58.520 78.038 13.549 SUMMARY time (in ms/op): (of 10 iterations) Operation Max Min Mean Std Dev --------- --- --- ---- ------- Directory creation 0.320 0.219 0.269 0.034 Directory stat 0.103 0.081 0.093 0.008 Directory rename 0.228 0.208 0.217 0.007 Directory removal 0.401 0.352 0.366 0.015 File creation 0.627 0.391 0.428 0.071 File stat 0.237 0.195 0.212 0.013 File read 0.430 0.372 0.396 0.016 File rename 0.273 0.201 0.222 0.021 File cross rename 32.581 30.299 31.654 0.675 File removal 0.401 0.356 0.381 0.016 Tree creation 0.586 0.032 0.244 0.189 Tree removal 0.017 0.010 0.013 0.002
2 MDT operation (with --max-inherit-rr=10):
mpirun --hostfile ./machine_file --mca routed direct --map-by node -np 16 ./src/mdtest -n 500 \
-i 5 -v --renameCrossDir --renameDestDir /mnt/lustre/rename_dst -P -u -d /mnt/lustre/mdtest
SUMMARY rate (in ops/sec): (of 10 iterations)
Operation Max Min Mean Std Dev
--------- --- --- ---- -------
Directory creation 5148.376 3382.648 4065.842 566.032
Directory stat 18719.181 14892.648 16491.074 1422.617
Directory rename 6556.759 4225.317 5150.864 848.020
Directory removal 4680.177 2757.504 3443.849 663.302
File creation 20141.997 18473.397 19488.327 565.500
File stat 31112.649 27544.882 29707.841 1151.295
File read 20757.651 18785.162 19748.132 797.785
File rename 40060.497 33879.232 36143.238 1777.142
File cross rename 555.181 489.168 534.888 20.116
File removal 20958.316 19236.684 20277.368 580.470
Tree creation 33.481 1.719 8.082 9.342
Tree removal 95.366 50.204 74.723 14.640
SUMMARY time (in ms/op): (of 10 iterations)
Operation Max Min Mean Std Dev
--------- --- --- ---- -------
Directory creation 2.365 1.554 2.000 0.257
Directory stat 0.537 0.427 0.488 0.042
Directory rename 1.893 1.220 1.589 0.244
Directory removal 2.901 1.709 2.390 0.388
File creation 0.433 0.397 0.411 0.012
File stat 0.290 0.257 0.270 0.011
File read 0.426 0.385 0.406 0.016
File rename 0.236 0.200 0.222 0.011
File cross rename 32.709 28.819 29.953 1.183
File removal 0.416 0.382 0.395 0.011
Tree creation 0.582 0.030 0.259 0.202
Tree removal 0.020 0.010 0.014 0.003
mpirun --hostfile ./machine_file --mca routed direct --map-by node -np 16 ./src/mdtest -n 500 \
-i 5 -v --renameCrossDir --renameDestDir /mnt/lustre/rename_dst -P -d /mnt/lustre/mdtest
SUMMARY rate (in ops/sec): (of 5 iterations)
Operation Max Min Mean Std Dev
--------- --- --- ---- -------
Directory creation 1068.986 966.217 1014.511 49.666
Directory stat 22026.015 18469.147 20050.129 1396.732
Directory rename 1810.878 1667.056 1739.928 52.968
Directory removal 1180.598 1092.735 1143.245 32.849
File creation 5093.213 4592.407 4825.595 241.412
File stat 42986.924 38724.233 41080.586 1723.832
File read 17318.122 15229.364 16056.340 837.172
File rename 5581.042 5176.476 5419.780 177.371
File cross rename 300.000 290.643 296.067 4.034
File removal 3303.411 3011.372 3146.996 122.785
Tree creation 111.881 81.733 89.341 12.704
Tree removal 280.106 139.239 235.866 57.592
SUMMARY time (in ms/op): (of 5 iterations)
Operation Max Min Mean Std Dev
--------- --- --- ---- -------
Directory creation 8.280 7.484 7.901 0.382
Directory stat 0.433 0.363 0.401 0.028
Directory rename 4.799 4.418 4.601 0.141
Directory removal 7.321 6.776 7.002 0.204
File creation 1.742 1.571 1.661 0.082
File stat 0.207 0.186 0.195 0.008
File read 0.525 0.462 0.499 0.025
File rename 1.545 1.433 1.477 0.049
File cross rename 55.050 53.333 54.050 0.740
File removal 2.657 2.422 2.545 0.099
Tree creation 0.012 0.009 0.011 0.001
Tree removal 0.007 0.004 0.005 0.002
Prototype with cross-MDT remote intent (directory per rank)
mpirun --hostfile ./machine_file --mca routed direct --map-by node -np 16 ./src/mdtest -n 500 \
-i 5 -v --renameCrossDir --renameDestDir /mnt/lustre/rename_dst -P -u -d /mnt/lustre/mdtest
SUMMARY rate (in ops/sec): (of 5 iterations)
Operation Max Min Mean Std Dev
--------- --- --- ---- -------
Directory creation 6903.193 5607.542 6170.828 624.057
Directory stat 19987.748 15444.822 17688.766 2034.637
Directory rename 8076.436 4496.870 5571.709 1464.479
Directory removal 3613.489 3052.994 3373.492 206.127
File creation 22617.265 20314.634 21415.285 961.808
File stat 35382.579 31824.847 33196.975 1358.646
File read 22647.491 18073.463 20501.750 1960.361
File rename 43154.328 39763.032 41087.617 1407.015
File cross rename 626.293 553.305 600.452 28.214
File removal 23695.846 18750.416 21411.747 1980.724
Tree creation 59.134 2.485 15.264 24.637
Tree removal 91.200 71.388 79.029 7.883
SUMMARY time (in ms/op): (of 5 iterations)
Operation Max Min Mean Std Dev
--------- --- --- ---- -------
Directory creation 1.427 1.159 1.307 0.128
Directory stat 0.518 0.400 0.457 0.053
Directory rename 1.779 0.991 1.502 0.319
Directory removal 2.620 2.214 2.379 0.151
File creation 0.394 0.354 0.374 0.017
File stat 0.251 0.226 0.241 0.010
File read 0.443 0.353 0.393 0.039
File rename 0.201 0.185 0.195 0.007
File cross rename 28.917 25.547 26.696 1.314
File removal 0.427 0.338 0.376 0.036
Tree creation 0.402 0.017 0.234 0.160
Tree removal 0.014 0.011 0.013 0.001
V-1: Entering PrintTimestamp...
mpirun --hostfile ./machine_file --mca routed direct --map-by node -np 16 ./src/mdtest -n 500 \
-i 5 -v --renameCrossDir --renameDestDir /mnt/lustre/rename_dst -P -u -d /mnt/lustre/mdtest
SUMMARY rate (in ops/sec): (of 5 iterations)
Operation Max Min Mean Std Dev
--------- --- --- ---- -------
Directory creation 2081.462 1868.925 1959.301 101.998
Directory stat 16827.007 13571.741 14500.075 1318.825
Directory rename 2072.866 1827.079 1922.569 97.052
Directory removal 1240.509 912.301 1072.732 124.817
File creation 4937.138 4657.256 4854.427 112.892
File stat 44902.208 38255.366 40419.230 2701.414
File read 16258.513 14440.411 15169.564 829.933
File rename 5472.177 5068.488 5251.367 171.430
File cross rename 433.315 378.367 401.872 23.175
File removal 3143.691 2779.128 2959.365 140.226
Tree creation 160.769 33.232 101.722 53.386
Tree removal 291.879 3.675 150.559 102.056
SUMMARY time (in ms/op): (of 5 iterations)
Operation Max Min Mean Std Dev
--------- --- --- ---- -------
Directory creation 4.281 3.843 4.092 0.209
Directory stat 0.589 0.475 0.555 0.045
Directory rename 4.379 3.859 4.169 0.205
Directory removal 8.769 6.449 7.540 0.887
File creation 1.718 1.620 1.649 0.039
File stat 0.209 0.178 0.199 0.013
File read 0.554 0.492 0.529 0.028
File rename 1.578 1.462 1.525 0.050
File cross rename 42.287 36.925 39.918 2.255
File removal 2.879 2.545 2.708 0.129
Tree creation 0.030 0.006 0.013 0.010
Tree removal 0.272 0.003 0.059 0.119