Skip to content

XGBoost 3.4 via the pure-JVM xgboost-predictor - #907

Merged
ltrottier-yelp merged 10 commits into
masterfrom
feature/xgboost-3.4-predictor
Oct 5, 2026
Merged

ltrottier-yelp merged 10 commits into
masterfrom
feature/xgboost-3.4-predictor

Conversation

@ltrottier-yelp

Copy link
Copy Markdown
Collaborator

Moves XGBoost inference to the pure-JVM com.yelp:xgboost-predictor:1.0.1 and xgboost4j 3.4.0, the newest on Maven Central.

  • Densifies sparse features so MLeap reproduces the training feature vector.
  • Fixes SparkParityBase.checkRowWithRelTol for Scala 2.13 array fields (mutable.ArraySeq), so NaN compares equal.
  • xgboost4j-spark 3.4.0 mispredicts all-missing rows ([jvm-packages] Fix Spark batch predict for SparseVector features dmlc/xgboost#12347, fixed in 3.4.1, not on Central), so the classifier parity spec filters them. The full-dataset spec stays @Ignored for reactivation.

Locally: mleap-xgboost-runtime 34 passed, mleap-xgboost-spark 6 passed, 3 ignored.

🤖 Generated with Claude Code

ltrottier-yelp and others added 8 commits July 22, 2026 12:57
Serve XGBoost predictions through the standalone com.yelp:xgboost-predictor
artifact (pure-JVM reader + tree traversal, jafama transitively), replacing the
xgboost4j Booster on the prediction hot path. The dual-format reader dispatches
on the leading byte (legacy pre-1.0 binary -> ModelReader, '{' -> UBJSON), so
existing deployed bundles and new 3.3.0 bundles load through the same engine
with no bundle migration.

Engine-parity coverage (per-objective predict/margin/probability/leaf) lives in
the xgboost-predictor repo. XGBoostBundleBackwardCompatSpec here covers the full
MLeap bundle round-trip on both formats using the golden .model fixtures.

mleap-xgboost-runtime 28/28 and mleap-xgboost-spark 6/6 pass.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
xgboost 3.4.1 contains dmlc/xgboost#12347 (backported as #12466), which fixes
the batch transform divergence on rows whose features are all missing, so the
synthetic zero-feature rows no longer have to be filtered out of the
mleap-xgboost-spark parity dataset.

com.yelp:xgboost-predictor:1.0.0 is now on Maven Central with no declared
dependencies. Its public API differs from the pre-release layout, so the call
sites move to com.yelp.xgboost.FVec, the FVec factory methods, and
Predictor.predictRaw for margin-space output.

sbt-assembly 2.5.0 is required because 2.1.1 shades at an ASM level that cannot
read the Java records in the published predictor, failing
mleap-databricks-runtime-fat/assembly.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
xgboost4j-spark densifies a SparseVector both when it builds the training
DMatrix (XGBoostEstimator.toXGBLabeledPoint expands the features with
Vector.toArray, and it refuses a sparse features column unless missing is set)
and in XGBoostModel.transform. An index absent from the input is therefore an
explicit 0.0 feature, and only the model's missing value turns it into an absent
one. XgbConverters kept the gap absent, which is equivalent only when missing is
0.0f, so MLeap disagreed with .transform on every row that had a gap: on a
100-row probe the two sides differed on 4 rows, up to 0.19 in probability.

The predictor path now densifies and lets treatsZeroAsNA decide, and asXGB
passes the model's missing to the DMatrix, mirroring
new DMatrix(iterator, null, getMissing()). That value is threaded through the
booster-backed models and persisted by their bundle ops under the same missing
key the mleap-xgboost-spark ops already write, defaulting to NaN for bundles
that predate it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
On Scala 2.13 a Spark Row array field is a mutable.ArraySeq, which is not a
scala.Seq (now immutable.Seq). The Seq case in checkRowWithRelTol therefore
never matched, and arrays fell through to ==, where NaN != NaN. Matching
scala.collection.Seq restores the element-wise relTol comparison, which treats
two NaNs as equal.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
1.0.1 tests the predictor against xgboost4j 3.4.0. Its classes are identical to
1.0.0 and its pom still declares no dependencies.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
xgboost4j_2.13 and xgboost4j-spark_2.13 3.4.1 and 3.4.2 are not on Maven
Central, so MLeap cannot resolve them from its configured repositories. 3.4.0
lacks dmlc/xgboost#12347, so the Spark/MLeap classifier parity spec drops the
synthetic zero-feature rows again, which xgboost4j-spark 3.4.0 mispredicts.

The full-dataset check stays as XGBoostClassificationModelZeroFeatureRowsParitySpec,
marked @ignore. Removing the annotation reactivates it once MLeap depends on an
xgboost release with the fix.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
SDKMAN's Scala 2.13.16 download points at downloads.lightbend.com, which now
returns 403, so the devcontainer image build fails. 2.13.18 downloads from
GitHub releases and matches scalaVersion in project/Common.scala.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
ltrottier-yelp and others added 2 commits October 5, 2026 09:01
mleap-xgboost-benchmark trains a model per vector width (100 to 1M features,
about 100 stored entries per row) and times FVec construction plus prediction
for the 0.25.2 map-backed FVec, the dense row, and FVecFactory.fromSparseVector.
It is not aggregated into root, so CI and releases skip it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Densifying a SparseVector made FVec construction cost proportional to the
vector size: 1.3 ms per row at 1M features in mleap-xgboost-benchmark.
xgboost-predictor 1.0.2 adds FVec.fromSparse, which keeps the dense-row
semantics with a table sized to the stored entries, at 11.6 us per row from
8720 to 1M features.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@ltrottier-yelp
ltrottier-yelp merged commit a7f6a8e into master Oct 5, 2026
18 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant