Skip to content

Latest commit

 

History

History
135 lines (101 loc) · 6.25 KB

File metadata and controls

135 lines (101 loc) · 6.25 KB

Internals

How NeRF.cpp works, and what it scores. For building and running it see the README and usage.md.

The training loop lives in src/main.cpp and is the best entry point: it is one function, and every stage below is a call from it.

File Contains
src/main.cpp training loop, losses, evaluation and preview scheduling
src/model.cpp positional encoding, SIREN and linear layers, the radiance field and proposal network
src/renderer.cpp ray generation, sampling strategies, volume rendering, sample_pdf, interlevel loss
src/data.cpp transforms.json parsing, image load and save
src/eval.cpp PSNR, RMSE, SSIM and the metrics CSV
src/utils.cpp device selection, orbit poses, preview renders, checkpoints
src/args.cpp the command-line flag table, which also generates --help

Training progression

The field converging during training. The object rotates while the reconstruction sharpens from the noisy SIREN initialisation to the final render.

Rays

For each pixel, NeRFRenderer::get_rays builds the camera-space direction ((x − W/2)/f, −(y − H/2)/f, −1), rotates it into world space with the frame's extrinsics, and takes the translation column as the ray origin.

For the proposal and stratified samplers the rays from every training image are concatenated once up front, and each step draws --ray-batch of them at random from that pool rather than training on one whole image at a time.

Sampling

Sample placement along a ray determines how much compute lands on surfaces rather than empty space. Three strategies:

  • uniform — evenly spaced points between near and far.
  • stratified — split the interval into bins, jitter one sample within each (NeRF §5.2). Better gradients than a fixed grid.
  • proposal — the default. A cheap coarse pass builds a per-ray density profile, and fine samples are drawn from it by inverse transform sampling, concentrating points near surfaces.

All three get the same main-network budget of 192 samples per ray, so comparisons are like for like. Single-pass samplers spend all 192 on one grid; the proposal path uses 64 coarse plus 128 importance samples, then evaluates the full radiance field at all 192 merged depths.

The coarse pass runs a separate density-only head (SirenNeRF::proposal_sigma): Fourier-encoded positions behind a stop-gradient, a small trunk, and a σ output. No view directions, no RGB. That is substantially cheaper than running the full model twice. An interlevel loss from Mip-NeRF 360 trains the coarse histogram to upper-bound the main network's weight distribution at different bin locations, as in the paper.

--warmup N makes the full model drive the coarse pass for the first N iterations, before handing over to the proposal head. The default of 0 uses the proposal head from step 0.

Network

Positions are Fourier-encoded at L = 10 and view directions at L = 4, both feeding SIREN trunks (src/model.cpp). Density is a function of position alone; colour also depends on view direction, which is what allows view-dependent effects such as specular highlights.

Component Width SIREN layers Head Output
Main trunk (trunk_) 128 6 (1 input + 5 hidden) — features
Main σ head 128 → 1 — linear + softplus density
Main RGB head 128 + view enc → 128 2 linear + sigmoid RGB
Proposal trunk (prop_trunk_) 128 3 (1 input + 2 hidden) — features
Proposal σ head 128 → 1 — linear + softplus density

The proposal network is a submodule of the main model, so model.parameters() covers both and one AdamW optimiser trains them together. It is trained only by the interlevel loss, never by the photometric loss.

Volume rendering

src/renderer.cpp turns σ along each ray into per-sample alpha and transmittance by discrete quadrature, then composites RGB and depth with the resulting weights over a configurable background colour (white by default).

Loss

Pseudo-Huber (Charbonnier) photometric loss by default, mean(sqrt(c² + e²) − c), plus the interlevel loss when training with the proposal sampler, and a small regulariser that keeps the first-layer SIREN frequencies from collapsing together.

One note on --huber-c, since it caused a real bug. An earlier default of 5.4e-4 * sqrt(3) sits below one 8-bit intensity level, so every pixel landed in the loss's linear region, gradients on flat areas never fell off, and backgrounds collapsed. The current 0.1 keeps ordinary errors in the quadratic basin. Under --loss mse the flag is ignored.

Evaluation

src/eval.cpp computes PSNR, RMSE and SSIM on the held-out views every --eval-every iterations (1000 by default). Evaluation renders use a deterministic sampler (uniform coarse grid, even-quantile importance) so the numbers are reproducible between runs.

Results

Trained on the NeRF synthetic lego and ship scenes at 160×160, 10,000 iterations, on an NVIDIA RTX A6000. Metrics are averaged over the 13 held-out views at the final evaluation step.

Scene PSNR ↑ SSIM ↑
lego 25.05 0.901
ship 25.03 0.767

Reproduce with defaults:

./build/NeRF.cpp data/lego output_lego_160_10k

Cost of the coarse pass

With hierarchical sampling, the coarse pass can either reuse the full radiance field or run the cheap proposal head. Timing one deterministic 120×120 test view, averaged over 10 runs on the same GPU with the lego model:

Coarse pass ms / view ↓
full model 167
proposal head 136

The proposal head is 19% faster because it skips view-dependent RGB and uses a narrow density-only trunk. The fine pass still costs a full 192-sample evaluation, so this is not a 19% end-to-end win; it makes the coarse PDF step cheap enough to run at both training and evaluation time.