A field report on the GPU-side cost of NVIDIA NVENC H.264 and AV1 encoding on an RTX 5090. The central question is simple: if NVENC is a dedicated hardware encoder, why does nvidia-smi report substantial Streaming Multiprocessor activity while ordinary H.264 and AV1 encoding is active?
Ranges are representative samples from nvidia-smi dmon -s u and HWMonitor screenshots, not laboratory averages. “Power” uses a clean live/current board-power sample when one was captured; contaminated session maxima are intentionally not treated as comparable measurements.
| ID | Application / stack | Driver / API path | Workload | Preview / source | SM | Memory | ENC | Board power | Notes |
|---|---|---|---|---|---|---|---|---|---|
| B0 | Baseline Bare desktop | 572.70-era baseline | No encode | OBS closed | 0% | — | 0% | — | System can genuinely settle at zero reported SM activity. |
| B1 | Baseline Firefox | 572.70-era baseline | No encode | Several active tabs | ~1% | — | 0% | — | Useful scale reference. |
| B2 | OBS OBS idle | 572.70 / OBS 32.2.2 | No encode | Preview off; Display Capture active | ~2–4% | ~2–3% | 0% | ~43.1 W | OBS itself has a persistent graphics cost. |
| B3 | OBS OBS idle | 572.70 / OBS 32.2.2 | No encode | Preview off; all sources removed | ~2% | — | 0% | — | Removing sources lowers but does not eliminate the OBS floor. |
| O1 | OBS H.264 | 572.70 / 13.0-era OBS | 2560×1440 @ 60, 12 Mbps | Preview off; Display Capture | ~5–11% | ~8–16% | ~11–21% | ~50.2 W | Original matched H.264 control. |
| O2 | OBS AV1 | 572.70 / 13.0-era OBS | 2560×1440 @ 60, 12 Mbps | Preview off; Display Capture | ~12–15% | ~15–16% | ~9–12% | ~58.2 W | AV1 was materially higher than H.264 in the original OBS comparison. |
| F1 | FFmpeg H.264 | 610.47 / NVENC API 13.1 | 2560×1440 @ 60, synthetic NV12, 12 Mbps | No OBS; null output | ~8–11% | ~8–9% | ~13–15% | Not cleanly captured | OBS, WebRTC, MediaMTX, capture and browser removed. |
| F2 | FFmpeg AV1 | 610.47 / NVENC API 13.1 | 2560×1440 @ 60, synthetic NV12, 8 Mbps | No OBS; null output | ~8–10% | ~5–6% | ~11–15% | Not cleanly captured | New API/client path largely equalized H.264 and AV1 SM activity. |
| O3 | OBS AV1 | 610.47 / OBS 32.2.2 13.0-era path | 2560×1440 @ 60, 8 Mbps | Preview off | ~13–16% | ~18–19% | ~10–13% | Not cleanly captured live | Driver update alone did not reduce OBS AV1 SM activity. |
| O4 | OBS AV1 | 610.47 / OBS 32.2.2 | 1920×1080 @ 60, 8 Mbps | Preview accidentally on | ~6–7% | ~13–14% | ~4–6% | Not cleanly captured live | Despite preview being enabled, substantially below 1440p60. |
| O5 | OBS AV1 | 610.47 / OBS 32.2.2 | 1280×720 @ 60, ~8 Mbps | Preview off | ~5% (very stable) | ~10% | ~1–2% | ~49.7 W live | 48-second run was unusually stable. |
| O6 | OBS AV1 | 610.47 / OBS 32.2.2 | 1920×1080 @ 120, ~8 Mbps | Preview off | ~9–10% | ~15–16% | ~6–8% | ~58.9 W live | Doubling FPS raises the encode-associated SM component substantially. |
OBS is not GPU-idle merely because the preview is disabled. OBS documents a dedicated graphics thread that renders the final mix as well as preview displays, converts the final texture into the configured output format, and then passes frames to encoders/outputs. This matches the observed ~2–3% SM floor with no active encode.
OBS backend design: the graphics pipeline has a dedicated graphics thread that renders the final mix; the final texture is converted to the configured backend video format before being sent onward.
This explains the OBS floor. It does not explain the standalone FFmpeg result, where OBS is absent and H.264/AV1 still produce ~8–10% SM activity.
| Metric | OBS idle | H.264 1440p60 | AV1 1440p60 |
|---|---|---|---|
| SM | ~2–4% | ~5–11% | ~12–15% |
| Memory | ~2–3% | ~8–16% | ~15–16% |
| Encoder engine | 0% | ~11–21% | ~9–12% |
| HWMonitor GPU utility | ~2% | ~6% | ~14% |
| HWMonitor Video Engine | 0% | ~39% | ~35% |
| Board power | ~43.1 W | ~50.2 W | ~58.2 W |
The original conclusion was obvious: under OBS 32.2.2 on 572.70, AV1 carried a much larger SM-side cost than H.264. The later FFmpeg/API-13.1 experiment changed that interpretation, but did not invalidate this measurement.
The current git-master Windows FFmpeg build refused to use NVENC on 572.70 because it required NVENC API 13.1; the dev partition was upgraded to 610.47. The synthetic input was generated as NV12 and encoded to the null muxer, eliminating OBS, capture, WebRTC, MediaMTX and browser rendering from the experiment.
| Metric | H.264 | AV1 | Interpretation |
|---|---|---|---|
| SM | ~8–11% | ~8–10% | Broadly equal |
| Memory | ~8–9% | ~5–6% | Not equal, but both active |
| Encoder engine | ~13–15% | ~11–15% | Both clearly use NVENC hardware |
| Application path | CPU-generated NV12 → FFmpeg NVENC wrapper → NVENC → null output | No OBS compositor or capture path | |
Common options deliberately disabled the obvious CUDA-assisted quality features: single pass, look-ahead 0, spatial AQ off, temporal AQ off, weighted prediction off, B-frames 0, B-ref off, temporal filtering 0; AV1 split encode was also disabled.
These runs kept the application, codec, driver and broadly the encoder setup constant while changing resolution and/or frame rate. The 1080p60 run accidentally left OBS preview enabled, which biases that row upward rather than making it look cheaper.
| Resolution / FPS | Pixels / frame | Pixels / second | SM | Approx. excess above ~3% OBS floor | Memory | ENC | Live board power |
|---|---|---|---|---|---|---|---|
| 1280×720 @ 60 | 0.922 MP | 55.30 MP/s | ~5% | ~2% | ~10% | ~1–2% | ~49.7 W |
| 1920×1080 @ 60 | 2.074 MP | 124.42 MP/s | ~6–7% | ~3–4% | ~13–14% | ~4–6% | Not cleanly captured live |
| 1920×1080 @ 120 | 2.074 MP | 248.83 MP/s | ~9–10% | ~6–7% | ~15–16% | ~6–8% | ~58.9 W |
| 2560×1440 @ 60 | 3.686 MP | 221.18 MP/s | ~13–16% | ~10–13% | ~18–19% | ~10–13% | Not cleanly captured live |
1080p120 processes about 248.8 million pixels/second, roughly 12.5% more than 1440p60 at 221.2 million pixels/second. Yet 1080p120 measured only ~9–10% SM while 1440p60 measured ~13–16%. Resolution/dimensions, surface handling, encoder partitioning, or another path-dependent factor matters; this is not a simple linear per-pixel tax.
At fixed 1080p resolution, moving from 60 to 120 FPS increased total SM from ~6–7% to ~9–10%. After subtracting the ~3% OBS floor, the encode-associated component moved from roughly ~3–4% to ~6–7%, close to doubling.
Power is useful but harder to compare cleanly because clocks, P-states, the NVIDIA “Prefer maximum performance” setting, and HWMonitor's persistent Max column can contaminate a session. Only clean current/live samples are promoted below.
| Test | Board power | Status | Comment |
|---|---|---|---|
| OBS idle / 572.70 | ~43.1 W | Clean sampled current | Display Capture active, no stream. |
| OBS H.264 1440p60 / 572.70 | ~50.2 W | Clean sampled current | About +7 W over that OBS idle sample. |
| OBS AV1 1440p60 / 572.70 | ~58.2 W | Clean sampled current | About +15 W over OBS idle; ~8 W over H.264 in that epoch. |
| OBS AV1 720p60 / 610.47 | ~49.7 W | Live screenshot | Preview disabled. |
| OBS AV1 1080p120 / 610.47 | ~58.9 W | Live screenshot | Preview disabled. |
The later 610.47 screenshots also contain HWMonitor Max values, but those maxima span earlier activity in the same monitoring session and are not treated as controlled per-test measurements.
The argument here is not that the NVENC ASIC is fake. The encoder-engine counter is active and NVIDIA's dedicated block plainly exists. The problem is that NVIDIA has repeatedly described NVENC in language that strongly implies the graphics/CUDA side is left free, while NVIDIA's own nvidia-smi telemetry reports substantial SM activity during ordinary encode workloads.
NVIDIA's transition story: before Kepler, GPU video encoding used the CUDA-core array; Kepler introduced specialized H.264 NVENC circuitry. NVIDIA explicitly separated optional CUDA pre-processing from the actual H.264 encoding performed by NVENC.
NVIDIA described NVENC as dedicated H.264 hardware that “does not use the GPU's graphics engine” and leaves that engine available for other work.
NVIDIA described fully accelerated encoding as independent of graphics performance and said complete encode offload leaves graphics bandwidth available for game rendering.
NVIDIA stated that NVENC/NVDEC are separate from CUDA cores and can run encoding/decoding without slowing concurrent graphics or CUDA workloads.
NVIDIA described NVENC as an independent physical section of the GPU dedicated to encoding, and in its XSplit guide went so far as to say the GPU can operate normally while that region streams or records.
GeForce RTX streaming article ↗ · NVIDIA NVENC/XSplit guide ↗
NVIDIA still describes NVENC as fully hardware based and independent of graphics/CUDA cores, stating that with end-to-end encoding offloaded, those cores are free for other operations.
NVIDIA separately acknowledges that some encoder features internally use CUDA: two-pass high-quality rate control, look-ahead, adaptive quantization, weighted prediction, RGB input encoding, temporal filtering, and hierarchical B-frame reference mode. NVIDIA says the graphics/CUDA impact of those features is minimal.
NVENC API Programming Guide — Encoder Features using CUDA ↗
The standalone FFmpeg tests intentionally used NV12 and disabled the applicable obvious suspects: two-pass/multipass, look-ahead, AQ, weighted prediction, temporal filtering and B-frame reference behavior. Substantial SM activity remained for H.264 and AV1.
The NVENC engine is active, but so are the SMs. This was reproduced in OBS and in a stripped standalone FFmpeg path. The report does not claim that the codec algorithm itself has secretly moved onto SMs; it does show that the real encode path is not SM-free.
That floor is explainable from OBS's architecture and exists with preview disabled. It must not be misattributed to NVENC. The FFmpeg isolation test is important precisely because it removes OBS and still shows substantial SM activity.
On the original OBS/572.70 path, AV1 sat around ~12–15% SM versus H.264's ~5–11%. On 610.47 with current FFmpeg/API 13.1, H.264 and AV1 both landed around ~8–10%. Re-running OBS 32.2.2 on 610.47 kept AV1 around ~13–16%, so the driver alone did not remove the gap. Something about the newer client/API/configuration path materially changed the observed behavior.
Higher resolution raises the cost. Higher frame rate raises the cost. However, 1080p120 processes more pixels per second than 1440p60 and still uses markedly less SM. That points toward a resolution/path threshold, tiling/surface behavior, encoder partitioning, or another non-linear implementation detail.
NVIDIA has repeatedly said NVENC is independent of graphics/CUDA resources and, in several documents, used language implying graphics resources remain fully available. On RTX 5090, NVIDIA's own nvidia-smi reports meaningful SM activity during ordinary H.264 and AV1 encode paths even after documented CUDA-assisted encoder features are disabled. The mechanism is unresolved, but the simplistic marketing model is not an adequate description of the observed system behavior.
SM activity is not a direct reservation of 10% of shader throughput. Scheduling, occupancy, clocks, instruction mix and bottlenecks matter. Real contention still requires a GPU-bound performance test if the practical FPS cost needs to be quantified.
NVIDIA's Video Codec SDK 13.1 includes source samples and Windows build instructions. The official AppEncD3D11 sample submits ID3D11Texture2D resources to NVENC and is the closest available NVIDIA-owned path for removing FFmpeg and OBS from the client side.
-nv12 mode first converts BGRA to NV12 using D3D11/DXVA VideoProcessorBlt, then submits the NV12 texture to NVENC. That gives a useful NVIDIA-owned A/B path, but the conversion step is still another engine in the experiment.| Candidate | Advantage | Contamination | Usefulness |
|---|---|---|---|
AppEncD3D11 default RGB | Official NVIDIA sample; direct D3D11 texture path | RGB encode is explicitly documented as CUDA-assisted | Bad for testing “core NVENC only” |
AppEncD3D11 -nv12 | Official NVIDIA sample; NV12 submitted to NVENC | BGRA→NV12 conversion via DXVA VideoProcessBlt | Good H.264-vs-AV1 A/B; imperfect absolute baseline |
| Small modified AppEncD3D11 accepting native NV12 | Closest to D3D11 NV12 → NVENC only | Requires a small source modification / raw input path | Best final “purity” test if desired |
NVIDIA Video Codec SDK download ↗ · SDK 13.1 Read Me / build instructions ↗ · Official AppEncD3D11 source ↗
Tizlik.Telemetry.md remains the chronological lab notebook. obs-profiles.json contains the current Tizlik OBS recommendations. README.md documents the current Tizlik/MediaMTX implementation.