Bake It Till You Make It

Ultrafast Spatial Texture-Surfel Splatting
Preprint · 2026
1Technical University of Munich  ·  2Siemens Healthineers
Teaser: textured surfel novel view synthesis
Our textured surfel representation decouples high-frequency texture from view-dependent geometry and appearance, enabling low primitive counts and real-time rendering across desktop and mobile hardware.
2600+
FPS on every benchmark
(fixed-function Vulkan, RTX 5090)
12–20×
faster than 3DGS
at over 10× fewer primitives
No MLP
pure texture lookups
— real-time on a phone

Abstract

Recent extensions of 3D Gaussian Splatting capture fine color detail using hash-grid based appearance parameterization, but pay a high per-fragment cost from neural field queries during rendering. We introduce a decoupled radiance representation that models low-frequency geometry and view-dependent appearance with 2D surfels, while representing high-frequency texture as a view-independent residual. Because the residual is view independent, we bake it into a compact texture, so inference reduces to pure hardware texture sampling with no neural queries. Combined with sparsity-enhancing optimization, our method renders significantly faster and sparser than prior work while preserving perceptual fidelity, reaching real-time 4K rates on consumer hardware and real-time rendering on mobile devices.

Key Ideas

Method

Full pipeline: per-surfel view-dependent colour plus a view-independent texture residual, baked and compressed for inference
During training, view-dependent colour is modelled per surfel while the view-independent texture residual is learned separately. For inference the residual is baked and compressed, so rendering uses only hardware texture sampling — no neural queries per fragment.

Interactive Demo

Explore a baked scene in your browser. Runs on WebGPU with the same texture-only renderer.

Results

Rendering throughput and quality on three standard benchmarks, measured on a single RTX 5090. All timings use per-frame cuda.Event pairs (50 warmup + 400 timed frames, sort and preprocess included), cycling through each scene's full held-out test set so all methods are timed on identical images. We first highlight the two most relevant baselines: FastGS, the fastest untextured method, which optimises the per-primitive cost of a conventional splat, and Nexels (100K), the strongest textured baseline near our primitive budget. We instead change what a primitive stores — moving high-frequency appearance into a baked texture, so far fewer and cheaper primitives are needed. The full comparison follows below.

BenchmarkMethodPrimitivesFPS Our speed-up PSNR SSIM LPIPS 
Mip-NeRF 360
9 scenes
FastGS394,18912091.27×27.440.79670.2589
Nexels (100K)99,93410015.4×26.550.77640.2245
Ours123,051153827.230.78920.2114
  + HW rasterizer2621
Tanks & Temples
2 scenes
FastGS242,13312431.78×24.000.84040.2093
Nexels (100K)99,93218112.3×22.920.82440.1752
Ours86,746221724.100.85070.1447
  + HW rasterizer2998
Deep Blending
2 scenes
FastGS216,74814511.39×30.020.90450.2651
Nexels (100K)99,8438623.4×30.000.90750.2097
Ours66,622201429.870.89020.2176
  + HW rasterizer4347

Full comparison: one GPU, one eval script

Every baseline below is retrained and re-benchmarked by us on the same RTX 5090, with quality recomputed from the saved renders by one shared evaluation script — no numbers are copied from other papers. We are the fastest method on all three benchmarks; among methods above 1000 FPS we have the best LPIPS on every benchmark mean and on 12 of 13 scenes. Red, orange and yellow tints mark the best, second and third value in each quality and speed column (Primitives and Size are context, not ranked; see the note below).

LPIPS versus FPS on Mip-NeRF 360, Tanks and Temples and Deep Blending: our method sits alone at the fast end of the Pareto frontier on all three benchmarks
LPIPS vs FPS (LPIPS decreases upward, so the best corner is top right) for every method in the comparison plus variants. Solid red links are our quality/speed ladder (HQ → no reg → production); dashed links connect each configuration to the identical bake on the GPU's fixed-function rasterizer (hollow stars, measured quality parity). The frontier ends at our hardware points on every benchmark; on Tanks & Temples it is entirely ours.
PSNR versus FPS on the three benchmarks with the same methods, variants and hardware twins
The same comparison on PSNR (higher is better, axis at the bottom). Beta-Splatting anchors Mip-NeRF 360 and Tanks & Temples and Nexels-400K and FastGS (Big) hold slow steps on Deep Blending, but the fast end of every frontier is our hardware points.
SSIM versus FPS on the three benchmarks with the same methods, variants and hardware twins
And on SSIM, our weakest metric: dense baselines and FastGS (Big) hold more of the slow end, while the fastest frontier points remain ours on all three benchmarks.

Mip-NeRF 360 (9 scenes)

MethodPSNR SSIM LPIPS PrimsSize (MB)FPS
3DGS27.540.81900.21492,732,349646211
2DGS26.830.79890.25202,128,093495131
Beta-Splatting28.070.83110.19043,111,111356110
Beta-Splatting (200K)26.660.77250.2942205,33123469
BBSplat26.800.78740.2355237,77817648
Nexels (40K)25.990.76190.241139,977138119
Nexels (100K)26.550.77640.224599,934282100
Nexels (400K)27.210.80180.2057399,77123078
Speedy-Splat26.910.78830.2875315,961751259
FastGS27.440.79670.2589394,189981209
FastGS (Big)27.800.81910.21601,161,438890
Content-Aware Texturing26.840.78860.2340155,756436137
Textured Gaussians25.860.73360.2857100,0003,83728
NeST26.390.77340.2257961,16322424
Ours (HQ)27.880.80700.1981248,341839706
Ours (no reg)27.650.80150.2048168,017596962
Ours27.230.78920.2114123,0514451538
  + HW rasterizer2621

Tanks & Temples (2 scenes)

MethodPSNR SSIM LPIPS PrimsSize (MB)FPS
3DGS23.740.85740.16921,574,592372259
2DGS23.150.83670.2118851,362198230
Beta-Splatting24.710.87350.14341,750,000200190
Beta-Splatting (200K)23.380.83480.2144200,00023524
BBSplat23.700.85700.1501300,00022680
Nexels (40K)22.250.79420.211339,979138204
Nexels (100K)22.920.82440.175299,932282181
Nexels (400K)23.620.84490.1576399,767230137
Speedy-Splat23.440.82520.2399181,817431459
FastGS24.000.84040.2093242,133601243
FastGS (Big)24.360.85720.1758543,997959
Content-Aware Texturing23.370.83870.1973133,880127270
Textured Gaussians22.700.80580.2161100,0003,83730
NeST22.540.81780.1866378,4489048
Ours (HQ)24.500.85980.1392149,6815391161
Ours (no reg)24.360.85780.1415107,4283971524
Ours24.100.85070.144786,7463252217
  + HW rasterizer2998

Deep Blending (2 scenes)

MethodPSNR SSIM LPIPS PrimsSize (MB)FPS
3DGS29.710.91060.23752,471,124584216
2DGS29.470.90710.25661,508,720351157
Beta-Splatting29.440.90900.23723,000,000343136
Beta-Splatting (200K)29.640.90240.2756200,00023520
BBSplat29.340.90500.2570160,00011149
Nexels (40K)29.340.89810.230639,958138106
Nexels (100K)30.000.90750.209799,84328286
Nexels (400K)30.410.91380.2046399,18823056
Speedy-Splat29.620.90740.2678250,990591501
FastGS30.020.90450.2651216,748541451
FastGS (Big)30.310.91150.2437649,6031276
Content-Aware Texturing29.990.91410.2410185,094252147
Textured Gaussians29.040.89130.2652100,0003,83730
NeST28.860.90210.2273478,00011338
Ours (HQ)30.280.90000.2095133,095498961
Ours (no reg)30.250.90080.211793,8133661237
Ours29.870.89020.217666,6222592014
  + HW rasterizer4347

Size is the payload as each method stores it, so the column mixes compressed and uncompressed formats: our atlas ships BC7-compressed, while Content-Aware Texturing and Textured Gaussians store fp32 textures with no compressed variant in their pipelines. Beta-Splatting (200K) and Nexels (40K/100K) are those papers' own reduced primitive budgets, closest to our primitive count. NeST is run as its released baseline method. FastGS (Big) is FastGS's own train_big.sh recipe, its higher-quality/slower configuration; size is unrecorded because those checkpoints were not packaged for deployment. Ours (no reg) and Ours (HQ) ablate our two speed-critical design choices, each retrained, baked and finetuned identically. They form a ladder rather than independent knobs: the falloff regulariser acts on the Beta shape parameter, so the Gaussian variant (HQ) is necessarily unregularised. Adopting the compact Beta kernel is worth 1.3–1.4× the frame rate, switching the regulariser on a further 1.5–1.6×, and together 1.9–2.2×, for a combined 0.4–0.7 dB of PSNR. Both steps shrink the model (248K primitives, then 168K, then 123K on Mip-NeRF 360), and neither alternative keeps the throughput lead: at 962 and 706 FPS they fall below FastGS and FastGS (Big).

+ HW rasterizer rows draw the identical production bake with the GPU's fixed-function rasterizer and ROP blending (a Vulkan port of our renderer; Vulkan GPU timestamps, same 400-frame protocol). Quality parity is measured (≤0.001 LPIPS across all 13 scenes), so only speed changes. The same port helps dense baselines far less — 1.28× for both 3DGS and FastGS on garden, whose frames are dominated by sorting millions of primitives — while our ~100K-primitive scenes gain 1.7×: sparsity and compact kernels are what the hardware path rewards. On this path texturing costs only 3.9–8.6% of the frame, and our no-reg and HQ variants reach 1723–2945 and 1172–1932 FPS (dashed links in the figure).

Our method carries ~3× fewer primitives than FastGS on every benchmark (over 10× fewer than 3DGS) and is 7–9× faster than 3DGS on the benchmark means, up to 13× on garden. FastGS holds a small PSNR edge on Mip-NeRF 360 (−0.21 dB) and Deep Blending (−0.15 dB) and we lead on Tanks & Temples (+0.10 dB), while LPIPS favours us on all 13 scenes — the explicit texture holds high-frequency detail that a per-primitive colour model cannot. Nexels (400K), the strongest textured baseline, edges our LPIPS on Mip-NeRF 360 and Deep Blending but renders 16–36× slower and trails us on Tanks & Temples. FastGS's own high-quality recipe, FastGS (Big), closes much of the LPIPS gap and still trails us on every benchmark mean, at 1.6–2.3× our primitive count and 0.6–0.8× our frame rate.

Texture finetune. The baked atlas is initialised from the trained neural residual and then finetuned against the training images for a short schedule, with geometry and view-dependent colour frozen so only the texels move. Baking alone can only lose information relative to its teacher; optimising the texels directly recovers it and then exceeds it, worth +0.49 dB and −0.037 LPIPS on average over the raw bake at identical primitive count, storage and frame rate.

Protocol. FastGS checkpoints were trained with its own released script (train_base.sh), which loads images at the 3DGS default (long edge capped at 1600 px). On Mip-NeRF 360 we render them at the standard evaluation resolution (images_4 outdoors, images_2 indoors) so both methods are measured on identical images; retraining FastGS at that resolution changes its primitive count by −8 % and its throughput by <3 %. On Tanks & Temples and Deep Blending both methods already use the same native resolution. FastGS runs at its own released default (mult = 0.5); at mult = 1.0 it gains ~0.2 dB PSNR and loses ~11 % throughput, with LPIPS essentially unchanged.

BibTeX

@article{kelkar2026bake,
  title   = {Bake It Till You Make It: Ultrafast Spatial Texture-Surfel Splatting},
  author  = {Kelkar, Neel and Niedermayr, Simon and Petkov, Kaloian and
             Engel, Klaus and Westermann, R{\"u}diger},
  journal = {arXiv preprint arXiv:2607.13808},
  year    = {2026}
}