Papers
Topics
Authors
Recent
Gemini 2.5 Flash
Gemini 2.5 Flash
102 tokens/sec
GPT-4o
59 tokens/sec
Gemini 2.5 Pro Pro
43 tokens/sec
o3 Pro
6 tokens/sec
GPT-4.1 Pro
50 tokens/sec
DeepSeek R1 via Azure Pro
28 tokens/sec
2000 character limit reached

A Performance Evaluation of a Quantized Large Language Model on Various Smartphones (2312.12472v1)

Published 19 Dec 2023 in cs.LG, cs.AI, and cs.PF

Abstract: This paper explores the feasibility and performance of on-device LLM inference on various Apple iPhone models. Amidst the rapid evolution of generative AI, on-device LLMs offer solutions to privacy, security, and connectivity challenges inherent in cloud-based models. Leveraging existing literature on running multi-billion parameter LLMs on resource-limited devices, our study examines the thermal effects and interaction speeds of a high-performing LLM across different smartphone generations. We present real-world performance results, providing insights into on-device inference capabilities.

User Edit Pencil Streamline Icon: https://streamlinehq.com
Authors (6)
  1. Tolga Çöplü (3 papers)
  2. Marc Loedi (1 paper)
  3. Arto Bendiken (4 papers)
  4. Mykhailo Makohin (1 paper)
  5. Joshua J. Bouw (2 papers)
  6. Stephen Cobb (3 papers)
Citations (3)