$ cat /etc/cookies.conf
We use cookies to understand how people use this site.
Analytics cookies help us improve your experience.
They are off by default. Nothing tracks you until you say so.
$ select cookie_preferences
Members-Only
Recent Talks & Demos are for members only
You must be an AI Tinkerers active member to view these talks and demos.
Learn to accelerate and optimize AI/ML inference on LLMs using GPUs and batch processing, with potential vector database results. This talk covers production AI model powering.
Most production AI models using deep learning will be performing computational forward passes for inference. This demo will go into the steps you would take to accelerate and optimize inference on models like LLMs with the power of GPUs and batch processing. We may also be showing the results via a vector database.
Loading recent emails...