Skip to content

Latest commit

 

History

History
187 lines (149 loc) · 8.27 KB

File metadata and controls

187 lines (149 loc) · 8.27 KB

Usage

Running NeRF.cpp on your own data. For the build and a first run on the bundled lego scene, see the README. For how the algorithm works, see internals.md.

NeRF.cpp <data_path> <output_path> [options]

data_path is the directory containing transforms.json and the images it references. output_path is created if it does not exist.

Output files

File What it is
checkpoint.pt model weights, rewritten every --preview-every iterations; replayable with --render
frame_<iter>_<n>.png RGB preview renders during training
frame_depth_<iter>_<n>.png matching depth previews
frame_final_<n>.png a 30-view orbit rendered after training
eval_metrics.csv per-view PSNR/RMSE/SSIM, appended every --eval-every iterations
final_metrics.txt averaged metrics at the last evaluation

Every 8th view is held out as a test view and never trained on. For a 100-frame scene that is 13 test views and 87 training views.

scripts/make_gifs.sh assembles these frames into orbit GIFs and a training-progress GIF:

bash scripts/make_gifs.sh <output_dir> [n_final_frames] [max_iter] [plot_freq] [n_preview_frames] [dest_dir]

It needs ImageMagick, which the conda environment already provides. If you pass dest_dir, ffmpeg is used to produce smaller copies there.

If the orbit frames are gone but checkpoint.pt remains, set RENDER_DATA to the scene directory and the script re-renders them from the checkpoint before assembling:

RENDER_DATA=data/lego RENDER_SIZE=160 bash scripts/make_gifs.sh output_lego 30

RENDER_SIZE must match the --size the checkpoint was trained at, because --size determines the focal length. NERF_BIN overrides the binary path (default ./build/NeRF.cpp).

Render from a checkpoint

--render FILE loads weights and writes the --final-frames orbit without training:

./build/NeRF.cpp data/lego output_lego --render output_lego/checkpoint.pt --final-frames 30

The scene directory is still required, since the focal length comes from its transforms.json and --size. Pass the same --size, --samples and --importance used for training to reproduce that run's frames exactly; larger sampling values are valid too and simply render the same weights more finely.

Use your own scene

The loader reads a single transforms.json in your data directory, in the format used by the original NeRF synthetic dataset:

{
  "camera_angle_x": 0.6911112070083618,
  "frames": [
    {
      "file_path": "./train/r_0",
      "transform_matrix": [
        [-0.9999, 0.0042, -0.0133, -0.0538],
        [-0.0140, -0.2997, 0.9539, 3.8455],
        [0.0000, 0.9540, 0.2997, 1.2081],
        [0.0, 0.0, 0.0, 1.0]
      ]
    }
  ]
}

Constraints the loader imposes:

  • camera_angle_x is the horizontal field of view in radians, shared by all frames. Per-frame intrinsics are not supported.
  • file_path is relative to the directory holding transforms.json, and .png is appended to it. ./train/r_0 loads train/r_0.png.
  • transform_matrix is a 4×4 camera-to-world matrix in the OpenGL/Blender convention: camera looks down −Z, with +Y up. COLMAP emits the OpenCV convention (+Z forward, −Y up); convert by flipping the Y and Z axes, or the renders come out inverted.
  • Images are resized to a square --size × --size. Non-square input is stretched, so crop first.
  • The scene must lie between --near and --far. The defaults of 2 and 6 suit the synthetic dataset, where the object sits at the origin with cameras about 4 units out. Other scales need these adjusted; getting them wrong is the most common cause of a run converging to uniform fog.
  • A white background is composited where rays accumulate no density, matching the synthetic dataset's alpha-matted renders. Real captured backgrounds work, but the background is baked into the field.

For poses from your own captures, nerfstudio's ns-process-data wraps COLMAP and writes a transforms.json. You still need to convert the pose convention, and to supply a camera_angle_x if it emits per-frame intrinsics.

Flags

Every setting is a command-line flag, so changing one does not require a rebuild. --help prints the same list with current defaults.

Flag Default What it does
--render FILE none skip training: load this checkpoint and render the orbit
--iters N 10000 training iterations
--size N 160 images resized to N×N before training
--device auto|cpu|cuda auto auto falls back to CPU; cuda exits if no GPU is visible
--seed N 12345 RNG seed
--sampler proposal|stratified|uniform proposal see Sampling
--samples N 64 coarse samples per ray
--importance N 128 fine importance samples
--ray-batch N 8192 rays per optimisation step, pooled across all training images
--near F / --far F 2.0 / 6.0 the depth range sampled along each ray
--warmup N 0 iterations for which the full model, not the proposal head, drives the coarse pass
--interlevel-weight F 1.0 weight on the interlevel loss
--width N 128 trunk width
--depth N 5 hidden SIREN layers in the trunk
--lr F 5e-4 AdamW learning rate
--weight-decay F 1e-2 AdamW weight decay
--loss huber|mse huber photometric loss
--huber-c F 0.1 pseudo-Huber transition point
--batch-size N 1280000 sample points per forward chunk
--log-every N 50 iterations between loss lines
--eval-every N 1000 iterations between test-view evaluations
--preview-every N 100 iterations between preview renders and checkpoints
--preview-frames N 5 views per preview render
--final-frames N 30 views in the final orbit

Two common cases:

# Sanity check: does the data load and does the loss decrease?
./build/NeRF.cpp data/mine out --iters 200 --size 64 --ray-batch 1024

# Out of GPU memory. Lower --batch-size first, then --ray-batch, then --size.
./build/NeRF.cpp data/mine out --batch-size 320000 --ray-batch 4096

--batch-size controls only how many sample points go through the network per chunk. Lowering it costs throughput but leaves the result identical. --ray-batch and --size change what the model actually sees. Preview and final renders cap the chunk at 320000 regardless, since they push far more points per pass than a training step; lowering --batch-size below that lowers previews too.

Troubleshooting

Could not find a package configuration file provided by "Torch" — CMake cannot see LibTorch. With conda, the environment must be active in the shell you run cmake from (conda activate nerfcpp); CMake reads $CONDA_PREFIX. Otherwise unzip LibTorch to ./libtorch in the repo root, or configure with -DCMAKE_PREFIX_PATH=/path/to/libtorch.

Link errors mentioning __cxx11 or std::string ABI — a system compiler was used against conda's LibTorch. Build from inside the activated environment so the conda cxx-compiler is picked up, and delete build/ before reconfiguring.

CUDA::nvToolsExt target not found — already handled. CUDA 12 dropped the standalone NVTX target that LibTorch's CMake still references, so CMakeLists.txt defines an empty stub. Variants of this failure are worth tracing back to that block.

CUDA GPU is required for training, but LibTorch cannot see one — --device cuda was passed and no GPU was found. environment.yml installs a CUDA LibTorch by default, so this usually means nvidia-smi fails, or the driver is older than the CUDA 13 packages need — see the version note in that file. Dropping the flag falls back to CPU.

CUDA out of memory — lower --batch-size first; it affects throughput only, not results. Then --ray-batch, then --size.

Could not open '.../transforms.json' — the data path is the directory containing transforms.json, not the file itself.

Renders come out inverted — pose convention. See Use your own scene.

Training converges to uniform grey fog — usually --near/--far do not bracket the scene.