samuela Samuel Ainsworth
github.com/samuela · maintains 189 packages (92 directly), sole maintainer of 79, member of 1 teams
189 packages · page 1 of 2
| Package ▲ | Version | Maintainers | # | Teams | Deps | Used by | Used by (transitive) | Status |
|---|---|---|---|---|---|---|---|---|
Debug tools for AMD systems |
0.2.20 → 0.2.21 |
1 | 28 | 0 | 0 | single outdated | ||
C++ library for modeling and solving large, complicated optimization problems |
2.2.0 |
7 | 8 | 14 | 105 | maintained | ||
Experimental Rust-to-CUDA compiler that lets you write (SIMT) GPU kernels in safe(ish), idiomatic Rust |
0.2.1 |
6 | 9 | 0 | 0 | maintained | ||
Building blocks that make it easier to write safe and efficient CUDA C++ code |
12.9.27 |
6 | 6 | 29 | 117 | team-only unfree | ||
Samples for CUDA Developers which demonstrates features in CUDA Toolkit |
12.8 |
6 | 18 | 0 | 0 | maintained unfree | ||
CUDA Runtime |
12.9.79 |
6 | 8 | 90 | 116 | team-only unfree | ||
Extracts information from CUDA binary files (both standalone and those embedded in host binaries) and presents them in human readable format |
12.9.82 |
6 | 6 | 3 | 27 | team-only unfree | ||
C-based interface for creating profiling and tracing tools designed for CUDA applications |
12.9.79 |
6 | 6 | 16 | 42 | team-only unfree | ||
Decode low-level identifiers that have been mangled by CUDA C++ into user readable names |
12.9.82 |
6 | 6 | 3 | 24 | team-only unfree | ||
Pre-built applications which use CUDA |
12.9.79 |
6 | 11 | 0 | 0 | team-only unfree | ||
12.9.88 |
6 | 6 | 0 | 0 | team-only unfree | |||
NVIDIA tool for debugging CUDA applications on Linux and QNX systems |
12.9.79 |
6 | 10 | 1 | 22 | team-only unfree | ||
Nsight Eclipse Plugins Edition |
12.9.79 |
6 | 6 | 0 | 0 | team-only unfree | ||
CUDA compiler driver |
12.9.86 |
6 | 8 | 82 | 129 | team-only unfree | ||
Extracts information from standalone cubin files and presents them in human readable format |
12.9.88 |
6 | 6 | 3 | 27 | team-only unfree | ||
C-based programmatic interface for monitoring and managing various states within Data Center GPUs |
12.9.79 |
6 | 7 | 9 | 38 | team-only unfree | ||
Collect and view profiling data from the command-line |
12.9.79 |
6 | 7 | 0 | 0 | team-only unfree | ||
Prune host object files and libraries to only contain device code for the specified targets |
12.9.82 |
6 | 6 | 1 | 22 | team-only unfree | ||
Runtime compilation library for CUDA C++ |
12.9.86 |
6 | 7 | 27 | 97 | team-only unfree | ||
C-based Application Programming Interface (API) for annotating events, code ranges, and resources in your applications |
12.9.79 |
6 | 6 | 15 | 40 | team-only unfree | ||
Cross-platform performance profiling tool for optimizing CUDA C/C++ applications |
12.9.79 |
6 | 6 | 0 | 0 | team-only unfree | ||
Low-level API for heterogeneous computing that runs on CUDA-powered GPUs |
12.9.19 |
6 | 6 | 0 | 0 | team-only unfree | ||
API for profiling CUDA runtime |
12.9.79 |
6 | 6 | 19 | 49 | team-only unfree | ||
Enables the creation of sanitizing and tracing tools that target CUDA applications |
12.9.79 |
6 | 6 | 1 | 22 | team-only unfree | ||
Wrapper substituting the deprecated runfile-based CUDA installation |
12.9 |
6 | 23 | 5 | 21 | team-only unfree | ||
GPU-accelerated library of primitives for deep neural networks |
9.22.0.52 |
7 | 9 | 21 | 26 | maintained unfree | ||
Python and C++ Graph API with SOTA attention (SDPA / Flash Attention), MoE grouped GEMM fusions, and FP8/MXFP8 kernels for Hopper and Blackwell GPUs |
1.30.0 |
6 | 10 | 4 | 4 | maintained | ||
Set of high-performance libraries and tools for accelerating quantum computing simulations at both the circuit and device level by orders of magnitude |
25.09.0.7 |
6 | 11 | 0 | 0 | team-only unfree | ||
CUDA Templates for Linear Algebra Subroutines |
4.7.1 |
6 | 12 | 0 | 0 | maintained | ||
575.57.08 |
6 | 7 | 0 | 0 | team-only unfree | |||
Fast GPU memory copy library based on NVIDIA GPUDirect RDMA technology |
2.6 |
6 | 5 | 2 | 33 | maintained | ||
Service which supports GPU memory export and import (NVLink P2P) and shared memory operations across OS domains in an NVLink multi-node deployment |
575.57.08 |
6 | 7 | 0 | 0 | team-only unfree | ||
CUDA Basic Linear Algebra Subroutine library |
12.9.2.10 |
6 | 8 | 46 | 75 | team-only unfree | ||
High-performance, multi-process, GPU-accelerated library for distributed basic dense linear algebra |
0.8.1.2360 |
6 | 9 | 0 | 0 | team-only unfree | ||
Library of GPU-accelerated linear solvers with sparse matrices |
0.6.0.5 |
6 | 9 | 0 | 0 | team-only unfree | ||
High-performance FFT product CUDA library |
11.4.1.4 |
6 | 9 | 23 | 49 | team-only unfree | ||
Library to leverage GDS technology |
1.14.1.1 |
6 | 11 | 6 | 41 | team-only unfree | ||
Helper module for the cuBLASMp library that allows it to efficiently perform communications between different GPUs |
10.3.10.19 |
6 | 7 | 27 | 53 | team-only unfree | ||
Collection of dense and sparse direct linear solvers and Eigen solvers |
11.7.5.82 |
6 | 10 | 26 | 52 | team-only unfree | ||
High-performance, distributed-memory, GPU-accelerated library that provides tools for solving dense linear systems and eigenvalue problems |
0.8.0.3126 |
6 | 10 | 0 | 0 | team-only unfree | ||
GPU-accelerated basic linear algebra subroutines for sparse matrix computations for unstructured sparsity |
12.5.10.65 |
6 | 8 | 25 | 55 | team-only unfree | ||
High-performance CUDA library dedicated to general matrix-matrix operations in which at least one operand is a structured sparse matrix with 50% sparsity ratio |
0.8.1.1 |
7 | 7 | 6 | 11 | maintained unfree | ||
GPU-accelerated tensor linear algebra library for tensor contraction, reduction, and elementwise operations |
2.3.1.0 |
6 | 7 | 3 | 3 | team-only unfree | ||
Library of primitives for image and signal processing |
12.4.1.87 |
6 | 7 | 2 | 23 | team-only unfree | ||
C++ support for interfacing with the NVIDIA Performance Primitives (NPP) library |
0.10.0.0 |
6 | 7 | 0 | 0 | team-only unfree | ||
APIs which can be used at runtime to combine multiple CUDA objects into one CUDA fat binary (fatbin) |
12.9.82 |
6 | 7 | 0 | 0 | team-only unfree | ||
NVIDIA NSCQ (Node Scalable Coherency Quantum) API for NVSwitch communication |
575.57.08 |
6 | 6 | 0 | 0 | team-only unfree | ||
APIs which can be used at runtime to link together GPU device code |
12.9.86 |
6 | 7 | 15 | 66 | team-only unfree | ||
Provides high-performance, GPU accelerated JPEG decoding functionality for image formats commonly used in deep learning and hyperscale multimedia applications |
12.4.0.76 |
6 | 7 | 1 | 1 | team-only unfree | ||
Accelerates the decoding and encoding of JPEG2000 images on NVIDIA GPUs |
0.9.0.43 |
6 | 6 | 0 | 0 | team-only unfree | ||
NVIDIA NVSwitch Device Manager library |
575.57.08 |
6 | 6 | 0 | 0 | team-only unfree | ||
Parallel programming interface for NVIDIA GPUs based on OpenSHMEM |
3.6.5-0 |
6 | 22 | 5 | 10 | maintained unfree | ||
Accelerates TIFF encode/decode on NVIDIA GPUs |
0.5.1.75 |
6 | 6 | 0 | 0 | team-only unfree | ||
Multi-GPU and multi-node collective communication primitives for NVIDIA GPUs |
2.32.3-1 |
7 | 12 | 27 | 32 | maintained | ||
Tests to check both the performance and the correctness of NVIDIA NCCL operations |
2.20.0 |
7 | 7 | 0 | 0 | maintained | ||
Interactive profiler for CUDA and NVIDIA OptiX |
2025.2.1.3 |
6 | 12 | 0 | 0 | team-only unfree | ||
System-wide performance analysis and visualization tool |
2025.1.3.140 |
6 | 30 | 0 | 0 | team-only unfree | ||
High-speed data compression and decompression library optimized for NVIDIA GPUs |
5.0.0.6 |
6 | 6 | 0 | 0 | team-only unfree | ||
GPUDirect Storage kernel driver to read/write data from supported storage using cufile APIs |
2.25.7 |
6 | 6 | 0 | 0 | team-only unfree | ||
NVIDIA NVLSM (NVLink Subnet Manager) for 4th-generation NVSwitch management |
2025.03.1 |
6 | 6 | 0 | 0 | team-only unfree | ||
Simple (Single-precision AX Plus Y) FindCUDAToolkit.cmake example for testing cross-compilation |
0-unstable-2023-07-11 |
6 | 7 | 0 | 0 | team-only | ||
SDK that facilitates high-performance machine learning inference |
10.16.1.11 |
6 | 8 | 4 | 4 | team-only unfree | ||
Open Source Software (OSS) components of NVIDIA TensorRT |
10.16.1 |
6 | 7 | 0 | 0 | maintained | ||
Inference of Meta's LLaMA model (and others) in pure C/C++ |
0.6.0 → 1.0.0 |
10 | 11 | 4 | 43 | maintained outdated | ||
Inference of Meta's LLaMA model (and others) in pure C/C++, with CUDA support |
0.6.0 → 1.0.0 |
10 | 15 | 0 | 0 | maintained outdated | ||
Inference of Meta's LLaMA model (and others) in pure C/C++, with ROCm support |
0.6.0 → 1.0.0 |
10 | 13 | 0 | 0 | maintained outdated | ||
Inference of Meta's LLaMA model (and others) in pure C/C++, with Vulkan support |
0.6.0 → 1.0.0 |
10 | 14 | 2 | 2 | maintained outdated | ||
Programming language and developer toolchain for heterogeneous computing |
1.0.0 |
2 | 16 | 0 | 0 | maintained unfree | ||
Multi-Joint dynamics with Contact. A general purpose physics simulator |
3.15.0 |
3 | 9 | 2 | 23 | maintained | ||
Iteratively adjust text position in matplotlib plots to minimize overlaps |
1.4.0 |
1 | 20 | 0 | 0 | single | ||
Python package for handling annotated data matrices in memory and on disk |
0.13.3.post0 → 0.13.4 |
1 | 29 | 1 | 1 | single outdated | ||
Python library for implementing GraphQL servers using schema-first approach |
1.0.1 → 1.1.1 |
1 | 29 | 0 | 0 | single outdated | ||
Efficiently Composable Data Augmentation on the GPU with Jax |
0.3.3 → 0.4.1 |
1 | 19 | 0 | 0 | single outdated | ||
Automatically create Faiss knn indices with the most optimal similarity search parameters |
2.18.0 |
1 | 24 | 0 | 0 | single | ||
Library for reading, creating and manipulating PDF files |
2.1.25 |
1 | 23 | 0 | 0 | single | ||
Contrastive Language-Image Pretraining |
1.0-unstable-2022-07-27 |
1 | 20 | 0 | 0 | single | ||
CUDA Core Compute Libraries for Python |
1.2.1 |
6 | 29 | 0 | 0 | maintained broken | ||
NumPy-compatible matrix library accelerated by CUDA |
14.2.0 |
6 | 35 | 0 | 0 | maintained | ||
Deep learning optimization library that makes distributed training and inference easy, efficient, and effective |
0.19.6 → 0.19.7 |
7 | 34 | 0 | 0 | maintained outdated | ||
Python library for working with nested data structures |
0.1.10 |
2 | 23 | 10 | 54 | maintained | ||
Efficiently read embedding in streaming from any filesystem |
1.8.1 |
1 | 20 | 1 | 1 | single | ||
Fast array utilities |
1.4.1 → 1.5.1 |
1 | 26 | 1 | 1 | single outdated | ||
Fast read/write of AVRO files |
1.12.1 → 1.13.0 |
1 | 25 | 7 | 65 | single outdated | ||
FFCV: Fast Forward Computer Vision |
1.0.0 → 1.0.2 |
2 | 29 | 0 | 0 | maintained outdated | ||
Google Cloud Artifact Registry API client library |
1.22.0 → 1.23.0 |
1 | 26 | 1 | 53 | single outdated | ||
Derive PyPI package description from Python package docstring |
1.1.1 |
1 | 18 | 3 | 6 | single | ||
Hatchling plugin to create optional-dependencies pinned to minimum versions |
0.2.0 |
1 | 18 | 0 | 0 | single | ||
Python API and command line interface for HDFS |
2.7.3 |
1 | 21 | 0 | 0 | single | ||
In-place file processing |
1.0.2 |
1 | 18 | 0 | 0 | single | ||
Common/best-practice Invoke tasks and collections |
4.0.2 |
1 | 29 | 0 | 0 | single | ||
Source-built JAX frontend: differentiate, compile, and transform Numpy code |
0.11.2 |
6 | 28 | 45 | 80 | maintained | ||
JAX XLA PJRT Plugin for NVIDIA GPUs |
0.11.2 |
7 | 23 | 1 | 1 | maintained | ||
JAX Plugin for CUDA12 |
0.11.2 |
7 | 24 | 0 | 0 | maintained | ||
JAX XLA PJRT Plugin for NVIDIA GPUs |
0.11.2 |
6 | 23 | 1 | 1 | maintained | ||
JAX Plugin for CUDA13 |
0.11.2 |
6 | 24 | 0 | 0 | maintained | ||
Prebuilt jaxlib backend from PyPi |
0.11.2 |
1 | 19 | 24 | 81 | single | ||
Open-weight TTS model with 82 million parameters |
0-unstable-2025-06-16 → 0.9.4 |
1 | 26 | 0 | 0 | single outdated | ||
Peptide-MHC I binding affinity prediction |
2.3.13 |
1 | 31 | 0 | 0 | single | ||
Parsing MHC nomenclature in the wild |
3.64.4 |
1 | 20 | 1 | 1 | single | ||
Tiny library for coding with large language models |
0.3.3 |
1 | 19 | 0 | 0 | single |