Update to the newest nccl #1

arturnn · 2023-06-22T09:37:25Z

No description provided.

Signed-off-by: Jonas Zhou <[email protected]>

Fix hang in corner cases of alltoallv using point to point send/recv. Harmonize error messages. Fix missing NVTX section in the license. Update README.

Add support for CUDA graphs. Fuse BCM Gen4 switches to avoid suboptimal performance on some platforms. Issue NVIDIA#439. Fix bootstrap issue caused by connection reordering. Fix CPU locking block. Improve CollNet algorithm. Improve performance on DGX A100 for communicators with only one GPU per node.

Fix memory leaks. Fix crash in bootstrap error case. Fix Collnet clean-up issue. Make PCI switch vendor/device optional for XML injection. Add support for nvidia-peermem module.

Fix crash when setting NCCL_MAX_P2P_NCHANNELS below nchannels. Fix hang during sendrecv dynamic NVB connection establishment on cubemesh topologies. Add environment variable to only use SHARP on communicators beyond a given number of ranks. Add debug subsystem to trace memory allocations. Fix compilation with TRACE=1. (Issue NVIDIA#505)

Add support for bfloat16. Add ncclAvg reduction operation. Improve performance for aggregated operations. Improve performance for tree. Improve network error reporting. Add NCCL_NET parameter to force a specific network. Add NCCL_IB_QPS_PER_CONNECTION parameter to split IB traffic onto multiple queue pairs. Fix topology detection error in WSL2. Fix proxy memory elements affinity (improve alltoall performance). Fix graph search on cubemesh topologies. Fix hang in cubemesh during NVB connections.

ncclGroup's containing operations of mixed datatype, element, or collective would induce crash.

Add new API for creating a reduction operation which multiplies the input by a rank-specific scalar before doing an inter-rank summation (see: ncclRedOpCreatePreMulSum). Improve CollNet (SHARP) performance of ncclAllReduce when captured in a CUDA Graph via user buffer registration. Add environment variable NCCL_NET_PLUGIN="<suffix>" to allow user to choose among multiple NCCL net plugins by substituting into "libnccl-net-<suffix>.so". Fix memory leak of NVB connections. Fix topology detection of IB Virtual Functions (SR-IOV).

missing `#include <cstring>`.

I noticed when I enabled `NCCL_DEBUG_SUBSYS=ALLOC` that this function is called thousands of times, making the log output unintelligible. Fortunately, this function can be implemented without heap allocations.

Display hints of cause so that it would be easier for user to debug. Also change the error type from InternalError to InvalidUsage as most of time this is caused by a mismatch in collective size or env settings.

Disable NET transport for intra-node communication by setting the env to 1 It provides an option to error out instead of falling back to NET when superior intra-node transports (P2P and SHM) are unavailable

This reverts commit 445bc19.

First part on collective mismatch, second part on internal errors

Add network communication through another GPU connected with NVLink (PXN). Add aggregation of messages coming from different local GPUs through PXN and going to the same destination. Add new v5 plugin API with grouped receives and tags. Add compat for v4 plugins. Add naming of NCCL threads to help debugging. Fix NVLink detection and avoid data corruption when some NVLinks are down. Add support for Relaxed Ordering for IB. Add profiling and timing infrastructure.

reduce diagnostic noise for ThreadSanitizer. Fixes NVIDIA#649

Signed-off-by: Felix Abecassis <[email protected]>

For easier interpretation of debug messages like "connection closed by peer", "peer message truncated" and "peer collective mismatch"

Fix bug with CollNet Fix bug with zero-bytes send/recv operations Fix NCCL_PARAM implementation to avoid taking a lock on every call Fix bug when setting NCCL_IB_QPS_PER_CONNECTION to more than one. Improve error reporting for network errors.

Make sure make install also installs the static library. Fixes NVIDIA#662

Improve allreduce performance when we have more than one network interface per GPU and we need to use PXN to close rings. Add support for PCI Gen5 on 5.4 kernels. Fix crash when setting NCCL_SET_THREAD_NAME. Fix random crash in init due to uninitialized struct. Fix hang on cubemesh topologies. Add P2P_DIRECT_DISABLE parameter to disable direct access to pointers within a process.

Optimize CUDA graph launch; avoid launching a CPU callback for intra-node operations. Simplify kernel common code to improve the latency of send/recv operations. Strengthen CUDA streams semantics. Change NET API to v6, to add dmabuf support. Add ncclGetLastError() function. Add ncclRemoteError code and use it for remote network errors. Support the use of a different NCCL_NET parameter per communicator. Add support for SHM and P2P transfers using cudaMemcpy.

Summary: NCCL_DEBUG_FILE does not work properly since the recent v2.13.4 updates (NVIDIA#682) because it nows sets `ncclDebugLevel` after parse `NCCL_DEBUG_FILE`. This patch move parsing `tempNcclDebugLevel` before processing `NCCL_DEBUG_FILE` to ensure `NCCL_DEBUG_FILE` is parsed only when `NCCL_DEBUG > NCCL_LOG_VERSION` (same as previous behavior) Differential Revision: D38415208 fbshipit-source-id: 5689bbb798e73efb9e8594557666987f07e89a30

Add support for improved fault tolerance: non-blocking mode, new init function with config, and ncclCommFinalize function. Reintroduce collnet+chain algorithm, alongside collnet+direct. Add LL protocol for intra-node P2P (on by default) and network communication (off by default). Use network instead of shared memory when performance is better. Fix: wait for CUDA graph destroy before destroying comm with linked graph resources. Remove aggressive polling during enqueue. Fix DMABUF fallback on MOFED 5.4 and earlier.

…t file

Make sure all calls calling cudaMalloc (including devCommSetup) are called before the last bootstrapBarrier. That way, we avoid calls to cudaMalloc be blocked by a NCCL kernel launched on another GPU by another thread which completed init faster. Resolve NVIDIA#623.

Closes issue 658

Add support for H100 (sm90). Make sure NCCL kernel honor user stream priorities.

Fixes NVIDIA#726

Fix crash with CollnetChain on some node topologies Fix hang when interleaving the capture of different graphs Fix hang during init in multi-threaded mode Fix potential data corruption with LL128 protocol on unaligned buffers. Fix CPU usage during preconnect Fixes double-free in the error path for ncclCommInitAll Workaround hang on H100 with Ring/LL128 on 2 GPUs.

Also repurpose dummy plugin as example, including headers and compat layers from v6 to v2.

Add support for CUDA 12.0, drop Kepler (sm_35). Support for H100 features. Make socket code more robust and protected. Solves NVIDIA#555. Improve performance on large CUDA graphs, reducing dependencies. Reduce inter-socket bandwidth on AMD CPUs to favor better paths. Various fixes to ncclCommAbort. Make service thread polling resistant to EINTR. Compile with profiling API by default. Extend NVTX instrumentation with call arguments.

NCCL Net v4 supports a maximum handle size of 64 bytes whereas the ext-net example header files set it for NCCL Net v3. Since, `aws-ofi-nccl` plugin plans to follow the example header files, fix it here. Signed-off-by: Rashika Kheria <[email protected]>

Add support for 400Gbit NDR network adapters (CX7) Handle EINTR in socket poll() function Add NCCL_PROGRESS_APPENDOP_FREQ to control op append overhead Resource cleanup fixes Fix double free in case of init failure Fix crash in ncclCommAbort Revert AMD speed commit

Add new NVLS algorithm for allreduce using NVLink SHARP (intra-node only). Add new config options: cgaClusterSize, minCTAs, maxCTAs, netName. Enable LL128 when we use PXN to close rings. NVTX3 includes update. Fix crash when one CollNet (SHARP) rail fails to initialize.

Shutdown socket before close in ncclSocketClose()

Add support for IB SHARP to NVLS (NVLink SHARP algorithm). Add NVLS+Tree algorithm. Add support for memory management using cuMem* functions. Use all NICs for Send/Receive operations on systems with more than one NIC per GPU (NVIDIA#804). Add ncclCommSplit primitive, with resource sharing option in config. Fix alltoallv hang (NVIDIA#788) Increase number of channels on H100 when we're not limited by NVLink. Improve error reporting in case of IB failure, printing local and remote ID (NVIDIA#779). Add build option to allow compilation against RDMA includes instead of dynamically loading IB verbs symbols (NVIDIA#802). Fix context creation for progress thread (NVIDIA#803). NET/IB: add option to use multiple QPs in round-robin mode. Fix tree performance issue when NVB is disabled on HCM topologies.

Fix data corruption with Tree/LL128 on systems with 1GPU:1NIC. Fix hang with Collnet on bfloat16 on systems with less than one NIC per GPU. Fix long initialization time. Fix data corruption with Collnet when mixing multi-process and multi-GPU per process. Fix crash when shared memory creation fails. Fix Avg operation with Collnet/Chain. Fix performance of alltoall at scale with more than one NIC per GPU. Fix performance for DGX H800. Fix race condition in connection progress causing a crash. Fix network flush with Collnet. Fix performance of aggregated allGather/reduceScatter operations. Fix PXN operation when CUDA_VISIBLE_DEVICES is set. Fix NVTX3 compilation issues on Debian 10.

jonaszhou1 and others added 30 commits December 17, 2020 11:15

x86: Add CPU detection for Zhaoxin processors

3996562

Signed-off-by: Jonas Zhou <[email protected]>

2.8.4-1

911d61f

Fix hang in corner cases of alltoallv using point to point send/recv. Harmonize error messages. Fix missing NVTX section in the license. Update README.

2.9.8-1

ca8485b

Fix memory leaks. Fix crash in bootstrap error case. Fix Collnet clean-up issue. Make PCI switch vendor/device optional for XML injection. Add support for nvidia-peermem module.

Fix to NVIDIA#560

5f2f2f6

ncclGroup's containing operations of mixed datatype, element, or collective would induce crash.

Fix Collnet when GDR is disabled

4ec992f

Fix compilation failure in "src/enqueue.cc" on older GCC because of

30ca3fc

missing `#include <cstring>`.

Perform busIdToInt64 on the stack.

8cf7325

I noticed when I enabled `NCCL_DEBUG_SUBSYS=ALLOC` that this function is called thousands of times, making the log output unintelligible. Fortunately, this function can be implemented without heap allocations.

Improve warning message about truncated messages

f589932

Display hints of cause so that it would be easier for user to debug. Also change the error type from InternalError to InvalidUsage as most of time this is caused by a mismatch in collective size or env settings.

Add env NCCL_NET_DISABLE_INTRA

c88c9f8

Disable NET transport for intra-node communication by setting the env to 1 It provides an option to error out instead of falling back to NET when superior intra-node transports (P2P and SHM) are unavailable

Build fastsocket plugin from ext-net

c5790b3

remove unused basePath

445bc19

Revert "remove unused basePath"

cc78e9f

This reverts commit 445bc19.

Fix ext-net/google-fastsocket build

0144073

Split IB parameter sanity check into two parts

fbfb6ac

First part on collective mismatch, second part on internal errors

Add pthread_detach()'s for threads we never pthread_join(). Helps

44eb40d

reduce diagnostic noise for ThreadSanitizer. Fixes NVIDIA#649

Remove unnecessary newline in plugin logging

1c7c014

Signed-off-by: Felix Abecassis <[email protected]>

Fix typo in net_ib.cc

b895abc

Display host name instead of numeric IP when referring to a peer

1382a87

For easier interpretation of debug messages like "connection closed by peer", "peer message truncated" and "peer collective mismatch"

Merge branch 'master' into truncated_msg_warning

2dfd837

Fix merging error

2247152

2.12.10-1

353e8ba

Fix bug with CollNet Fix bug with zero-bytes send/recv operations Fix NCCL_PARAM implementation to avoid taking a lock on every call Fix bug when setting NCCL_IB_QPS_PER_CONNECTION to more than one. Improve error reporting for network errors.

Merge remote-tracking branch 'origin/master'

8133784

Update Makefile to install static library.

9bfc1c6

Make sure make install also installs the static library. Fixes NVIDIA#662

kingchc and others added 23 commits August 18, 2022 11:50

Fix intermittent 11.6 builds: generate unique .cu file for each objec…

79fb032

…t file

address review comments

f89fd47

Use compatibility shim only with static cudart

78313a6

Closes issue 658

Merge remote-tracking branch 'origin/master'

99c28f2

2.15.1-1

da8152e

Add support for H100 (sm90). Make sure NCCL kernel honor user stream priorities.

Fixes a double-free in the error path of ncclCommInitAll.

2401f4a

Fixes NVIDIA#726

Merge tag 'v2.15.1-1'

d128d62

Merge tag 'v2.15.5-1'

2f4cb87

Add documentation for NCCL NET plugins

55b1d8a

Also repurpose dummy plugin as example, including headers and compat layers from v6 to v2.

Fix google-fastsocket plugin build

614b49f

2.16.5-1

f3d5166

Add support for 400Gbit NDR network adapters (CX7) Handle EINTR in socket poll() function Add NCCL_PROGRESS_APPENDOP_FREQ to control op append overhead Resource cleanup fixes Fix double free in case of init failure Fix crash in ncclCommAbort Revert AMD speed commit

2.17.1-1

5d3ab08

Add new NVLS algorithm for allreduce using NVLink SHARP (intra-node only). Add new config options: cgaClusterSize, minCTAs, maxCTAs, netName. Enable LL128 when we use PXN to close rings. NVTX3 includes update. Fix crash when one CollNet (SHARP) rail fails to initialize.

Shutdown socket before close in ncclSocketClose()

367e9b6

Add a comment to shutdown() in ncclSocketClose

006b6bc

Merge pull request NVIDIA#822 from KaimingOuyang/github/pytorch-hang-fix

9b7d5ed

Shutdown socket before close in ncclSocketClose()

arturnn mentioned this pull request Jun 22, 2023

Training fails on Vertex AI (GCP) due to NCCL error on A100 GPUs marian-nmt/marian-dev#998

Open

arturnn closed this Nov 12, 2024

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Update to the newest nccl #1

Update to the newest nccl #1

arturnn commented Jun 22, 2023

Update to the newest nccl #1

Update to the newest nccl #1

Conversation

arturnn commented Jun 22, 2023