XGBoost 3.4 via the pure-JVM xgboost-predictor - #907
Merged
Merged
Conversation
Serve XGBoost predictions through the standalone com.yelp:xgboost-predictor
artifact (pure-JVM reader + tree traversal, jafama transitively), replacing the
xgboost4j Booster on the prediction hot path. The dual-format reader dispatches
on the leading byte (legacy pre-1.0 binary -> ModelReader, '{' -> UBJSON), so
existing deployed bundles and new 3.3.0 bundles load through the same engine
with no bundle migration.
Engine-parity coverage (per-objective predict/margin/probability/leaf) lives in
the xgboost-predictor repo. XGBoostBundleBackwardCompatSpec here covers the full
MLeap bundle round-trip on both formats using the golden .model fixtures.
mleap-xgboost-runtime 28/28 and mleap-xgboost-spark 6/6 pass.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
xgboost 3.4.1 contains dmlc/xgboost#12347 (backported as #12466), which fixes the batch transform divergence on rows whose features are all missing, so the synthetic zero-feature rows no longer have to be filtered out of the mleap-xgboost-spark parity dataset. com.yelp:xgboost-predictor:1.0.0 is now on Maven Central with no declared dependencies. Its public API differs from the pre-release layout, so the call sites move to com.yelp.xgboost.FVec, the FVec factory methods, and Predictor.predictRaw for margin-space output. sbt-assembly 2.5.0 is required because 2.1.1 shades at an ASM level that cannot read the Java records in the published predictor, failing mleap-databricks-runtime-fat/assembly. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
xgboost4j-spark densifies a SparseVector both when it builds the training DMatrix (XGBoostEstimator.toXGBLabeledPoint expands the features with Vector.toArray, and it refuses a sparse features column unless missing is set) and in XGBoostModel.transform. An index absent from the input is therefore an explicit 0.0 feature, and only the model's missing value turns it into an absent one. XgbConverters kept the gap absent, which is equivalent only when missing is 0.0f, so MLeap disagreed with .transform on every row that had a gap: on a 100-row probe the two sides differed on 4 rows, up to 0.19 in probability. The predictor path now densifies and lets treatsZeroAsNA decide, and asXGB passes the model's missing to the DMatrix, mirroring new DMatrix(iterator, null, getMissing()). That value is threaded through the booster-backed models and persisted by their bundle ops under the same missing key the mleap-xgboost-spark ops already write, defaulting to NaN for bundles that predate it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
On Scala 2.13 a Spark Row array field is a mutable.ArraySeq, which is not a scala.Seq (now immutable.Seq). The Seq case in checkRowWithRelTol therefore never matched, and arrays fell through to ==, where NaN != NaN. Matching scala.collection.Seq restores the element-wise relTol comparison, which treats two NaNs as equal. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
1.0.1 tests the predictor against xgboost4j 3.4.0. Its classes are identical to 1.0.0 and its pom still declares no dependencies. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
xgboost4j_2.13 and xgboost4j-spark_2.13 3.4.1 and 3.4.2 are not on Maven Central, so MLeap cannot resolve them from its configured repositories. 3.4.0 lacks dmlc/xgboost#12347, so the Spark/MLeap classifier parity spec drops the synthetic zero-feature rows again, which xgboost4j-spark 3.4.0 mispredicts. The full-dataset check stays as XGBoostClassificationModelZeroFeatureRowsParitySpec, marked @ignore. Removing the annotation reactivates it once MLeap depends on an xgboost release with the fix. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
SDKMAN's Scala 2.13.16 download points at downloads.lightbend.com, which now returns 403, so the devcontainer image build fails. 2.13.18 downloads from GitHub releases and matches scalaVersion in project/Common.scala. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
mleap-xgboost-benchmark trains a model per vector width (100 to 1M features, about 100 stored entries per row) and times FVec construction plus prediction for the 0.25.2 map-backed FVec, the dense row, and FVecFactory.fromSparseVector. It is not aggregated into root, so CI and releases skip it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Densifying a SparseVector made FVec construction cost proportional to the vector size: 1.3 ms per row at 1M features in mleap-xgboost-benchmark. xgboost-predictor 1.0.2 adds FVec.fromSparse, which keeps the dense-row semantics with a table sized to the stored entries, at 11.6 us per row from 8720 to 1M features. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Moves XGBoost inference to the pure-JVM
com.yelp:xgboost-predictor:1.0.1and xgboost4j3.4.0, the newest on Maven Central.SparkParityBase.checkRowWithRelTolfor Scala 2.13 array fields (mutable.ArraySeq), soNaNcompares equal.@Ignored for reactivation.Locally:
mleap-xgboost-runtime34 passed,mleap-xgboost-spark6 passed, 3 ignored.🤖 Generated with Claude Code