We develop efficient tensor libraries for tensor computing, including tensor decompositions and tensor networks. We provide efficient tensor primitives on GPU tensor cores and Tensor Processing Units (TPUs). We optimize the data transfer, memory access, and support key tensor operations. E.g., cuTensor-tubal library fully exploits the separability in the frequency domain and maps the tube-wise and slice-wise parallelisms onto the single instruction multiple thread (SIMT) GPU architecture.
High-Performance Tensor Computing
GPU-Accelerated Tensor Computing Primitives for Machine Learning
COMPUTING THIRD-ORDER TENSORS ON GPUS
DESIGN OF THE CUTENSOR-TUBAL LIBRARY
OVERVIEW OF THE LIBRARY
PERFORMANCE EVALUATION
RELATED WORKS
CONCLUSION AND FUTURE WORK
TensorLet Team
Related Publications
[Book Chapter] X.-Y. Liu, Y. Fang, L. Yang, Z. Li, A. Walid. High-performance Tensor Decompositions for Compressing and Accelerating Deep Neural Networks. Tensors for Data Processing: Theory, Methods, and Applications. [Link] Elsevier; 2021 Nov 10.
[To Appear] N. Li, X. Zhi, W. Tong, and X.-Y. Liu. zkFinGPT: GPU-Accelerated Zero-Knowledge Proofs for Financial Generative Pre-trained Transformers. 2026. [NeurIPS Workshop] X.-Y. Liu, N. Li, K. Wang, X. Zhi, and W. Tong. zkFinGPT: Zero-Knowledge Proofs for Financial Generative Pre-trained Transformers. NeurIPS Workshop on Generative AI in Finance, 2025.
[ICDCS] X.-Y. Liu, J. Zhang, G. Wang, Weiqin Tong, and Anwar Walid. Efficient Pretraining and Finetuning of Quantized LLMs with Low-rank Structure. IEEE ICDCS, 2024. [arXiv version] FinGPT-HPC: Efficient Pretraining and Finetuning Large Language Models for Financial Applications with High-Performance Computing, Link.
[NeurIPS] X.-Y. Liu, Z. Zhang. Classical simulation of quantum circuits using reinforcement learning: parallel environments and benchmark. NeurIPS, Special Track on Datasets and Benchmarks, 2023.
[TC] X.-Y. Liu, H. Hong, Z. Zhang, W. Tong, J. Kossaifi, X. Wang, and A. Walid. High-performance Tensor-Train Primitives Using GPU Tensor Cores. IEEE Transactions on Computers, 2024.
[TC] X.-Y. Liu, Z. Zhang, Z. Wang, H. Lu, X. Wang*, and A. Walid. High-performance tensor learning primitives using GPU tensor cores. IEEE Transactions on Computers, 2022.
