Skip to content

Reducing Training Time and Enhancing Efficiency in Large Language Models: Ternary Output Quantization(TOQ)

Do-hoon Lee Primary Contact
Abstract

This study examines the background of recent advancements in artificial intelligence technology and the need for large-scale language models (LLM) and multimodal models (LMM), proposing the Ternary Output Quantization (TOQ) technique to reduce computational complexity and shorten training times for these models. TOQ simplifies the final outputs of models to three discrete values (-1, 0, +1), significantly reducing data processing requirements and enhancing the efficiency of model training and inference processes. Experimental results show that applying TOQ to BERT and DistilBERT models significantly reduces training times and improves both training and validation loss compared to traditional models. This paper presents the potential for wide application of TOQ across various forms of AI models, anticipating that it will facilitate easier training and deployment of models, especially in resource-constrained environments.

References
  1. Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L. and Polosukhin, I. (2017), Attention is All You Need, Advances in Neural Information Processing Systems, 30.
  2. Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., Hesse, C., Chen, M., Sigler, E., Litwin, M., Gray, S., Chess, B., Clark, J., Berner, C., McCandlish, S., Radford, A., Sutskever, I. and Amodei, D. (2020), Language models are few-shot learners, arXiv preprint arXiv:2005.14165.
  3. Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J. and Amodei, D. (2020), Scaling laws for neural language models, arXiv preprint arXiv:2001.08361.
  4. Narayanan, D., Shoeybi, M., Casper, J., LeGresley, P., Patwary, M., Korthikanty, V., Vainbrand, D., Kashinkunti, P., Bernauer, J., Catanzaro, B., Phanishayee, A. and Zaharia, M. (2021), Efficient large-scale language model training on GPU clusters, arXiv preprint arXiv:2104.04473.
  5. Hu, E., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L. and Chen, W. (2021), LoRA: Low-Rank Adaptation of Large Language Models, arXiv preprint arXiv:2106.09685.
  6. Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W., Rocktäschel, T., Riedel, S. and Kiela, D., (2020), Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks, arXiv preprint arXiv:2005.11401.
  7. Devlin, J., Chang, M.-W., Lee, K. and Toutanova, K. (2018), BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding, arXiv preprint arXiv:1810.04805
  8. Sanh, V., Debut, L., Chaumond, J. and Wolf, T. (2019), DistilBERT: A distilled version of BERT: smaller, faster, cheaper and lighter, arXiv preprint arXiv:1910.01108
  9. Mellempudi, N., Kundu, A., Mudigere, D., Das, D., Kaul, B. and Dubey, P. (2017), “Ternary Neural Networks with Fine-Grained Quantization”, arXiv preprint arXiv:1705.01462v3
  10. Zhu, C., Han, S., Mao, H. and Dally, W. J. (2016), “TRAINED TERNARY QUANTIZATION”, arXiv preprint arXiv:1612.01064v3.
  11. Joshua K. Cage, 임선집, 채호창(2023), “101가지 문제로 배우는 딥러닝 허깅페이스 트랜스포머 with 파이토치,” 32-90.
Keywords
AI Ternary Output Quantization LLM Learning Efficiency Data Processing Minimization
Details

Authors
Do-hoon Lee Primary Contact

  • Enrolled in the Master’s in AI•Big Data program at the Graduate School of AI, Seoul School of Integrated Sciences & Technologies (aSSIST)
  • Currently working at the Youth Counseling and Welfare Center, Gyeongsangnam-do Youth Support Foundation
  • Areas of interest : AI, Data Analysis etc.
How to Cite
Reducing Training Time and Enhancing Efficiency in Large Language Models: Ternary Output Quantization(TOQ). (2024). AI Journal of BUsiness, 1(1), 52-59. https://jnl.ampla.page/aijb/article/view/113