CUDA & Compute

CUDA, creator workloads, AI frameworks, and encoder/decoder blocks.

  • 17 Tracked terms
  • Last 30 days Feed window

What this topic collects on

An article joins this feed when it matches these terms. Each one is also a search of its own.

Latest in CUDA & Compute


dev.to > gde > gemma-4-on-a-tesla-t4-qat-weights-decode-179x-faster-than-bf16-2fi4

Gemma 4 on a Tesla T4: QAT Weights Decode 1.79x Faster Than bf16

1+ day, 23+ hour ago   (1298+ words) The GPU is already there. A T4 attached to a Compute Engine VM needs no queued resource, no instance launch and no image, so this rig has no provisioning tools at all. Everything it ships is about the software on the…...


dev.to > gde > serving-gemma-4-on-an-amd-mi300x-what-199-an-hour-buys-52h9

Serving Gemma 4 on an AMD MI300X: What $1.99 an Hour Buys

2+ day, 23+ hour ago   (1104+ words) The card is a DigitalOcean GPU droplet reached through AMD Developer Cloud (devcloud.amd.com) — same v2 API, same droplet ids, token from the My AMD Team account. Creating and destroying it are console actions, deliberately: both are dollar-per-hour decisions and…...


dev.to > junsik_yoo > a-pedagogical-introduction-to-porting-a-conjugate-gradient-solver-to-cuda-aih

A Pedagogical Introduction to Porting a Conjugate Gradient Solver to CUDA

4+ day, 33+ min ago   (1676+ words) GPU programming tutorials often begin with isolated examples: vector addition, reductions, matrix multiplication, and memory coalescing. These are useful for learning individual CUDA concepts, but there is another question that quickly arises when working with real scientific software: How do…...


dev.to > gde > a-4-gb-laptop-gpu-beats-a-12-core-cpu-by-43x-on-gemma-4-4150

A 4 GB Laptop GPU Beats a 12-Core CPU by 4.3x on Gemma 4

4+ day, 1+ hour ago   (736+ words) This article compares two ways of serving the same small language model on the same laptop: CPU-only, and on the 4 GB GTX 1650 Ti sitting in the same chassis. The payload is byte-identical on both arms and the command lines differ…...


dev.to > rtagl > -how-to-use-libvmafcuda-on-windows-an-easy-to-follow-guide-wsl2-docker-nvidia-59mc

How to use libvmaf_cuda on Windows: an easy-to-follow guide (WSL2 + Docker + NVIDIA)

4+ day, 1+ hour ago   (450+ words) TL;DR: Running GPU-accelerated VMAF on Windows is surprisingly painful. After spending hours fighting... Tagged with ffmpeg, nvidia, docker, tutorial....


mdpi.com > 2076-16/18/3417 > 9154

Applied Sciences, Vol. 16, Pages 9154: High-Throughput GPU Acceleration of MS-OSM for Multiple DM Trials with Optimized Kernel and Asynchronous I/O

5+ day, 8+ hour ago   (375+ words) Pulsar and fast radio burst observations widely adopt coherent dedispersion to compensate for dispersion effects introduced by the interstellar medium. The multisegment overlap-save method (MS-OSM) was proposed to alleviate the extremely large FFT requirement of conventional overlap-save (OSM) coherent dedispersion....


shattered.io > nvidia-cuda-toolkit-13-4-windows-on-arm-2026

NVIDIA CUDA Toolkit 13.4 Adds Windows on Arm Support

1+ week, 4+ hour ago   (556+ words) The release notes for CUDA NVCC 13.4.59 list supported architectures as x86_64, arm64-sbsa, and arm64 (Windows), spanning both Linux and Windows. That’s a meaningful detail for anyone tracking the toolkit’s evolution: arm64-sbsa (Server Base System Architecture) has been NVIDIA’s Linux-side Arm server target…...


dev.to > gde > gemma-4-on-an-old-4-gb-laptop-gpu-qat-takes-it-from-95-gib-to-16-b5l

Gemma 4 on an Old 4 GB Laptop GPU: QAT Takes It From 9.5 GiB to 1.6

1+ week, 3+ day ago   (1449+ words) This article provides a step by step deployment guide for Gemma 4 E2B's quantization-aware-trained (QAT) checkpoint to a local, laptop hosted GPU enabled system — a much older Lenovo Yoga 9 with a 4 GB GTX 1650 Ti. A suite of Python MCP tools is…...


etmm-online.com-online.com

Solidcam 2026 Boosts CAM Simulation With Machine Works GPU

1+ week, 6+ day ago   (146+ words) GPU-based simulation allows users to review large machining jobs more quickly. Solidcam 2026 integrates Machine Works GPU to generate in-process stock models in seconds, helping programmers check selected stages before running a full simulation. For the majority of machining operations Machine…...


daytona.io > changelog > mi355x-gpu-type-and-python-sdk-s3-upload-reliability

MI355X GPU type and Python SDK S3 upload reliability

2+ week, 3+ day ago   (48+ words) Daytona 0.210.0 adds a new GPU type to the API client and hardens S3 uploads in the Python SDK. API client: add MI355X GPU type Python SDK: improve S3 upload reliability sync go.sum for v0.207.1 Go SDK: bump to v0.210.0 New Partner with us © 2026 Daytona…...