Files
myWiki/raw/papers/Levada-CuBAS-curvature-adaptive-sampling-2026.md
2026-07-20 14:14:55 +08:00

2.3 KiB

CuBAS: Curvature-Based Adaptive Sampling for Supervised Classification (Raw Archive)

  • arXiv: 2607.03145v1
  • Author: Alexandre L. M. Levada (Federal University of São Carlos)
  • Published: July 3, 2026
  • Category: cs.LG, cs.AI, cs.CV, cs.IT, stat.ML
  • Code: https://github.com/alexandrelevada/CuBAS

Abstract

The informativeness of a training set is as consequential as its size, yet most sampling strategies remain agnostic to the intrinsic geometry of the data distribution. We introduce CuBAS (Curvature-Based Adaptive Sampling), an information-geometric framework for adaptive data selection in supervised classification, grounded in the q-state Potts Markov random field (MRF) model. The central insight is that a labeled dataset can be viewed as a statistical manifold, on which local curvature, estimated via the ratio of second- to first-order observed Fisher information, faithfully encodes the geometric complexity of the underlying data distribution. We construct a k-nearest-neighbor graph over the labeled data and derive a closed-form curvature score at each vertex from the Potts sufficient statistics. This curvature signal partitions the graph into two complementary regimes: low-curvature regions (smooth, homogeneous clusters) and high-curvature regions (decision boundaries, disproportionately informative). By selecting nodes from both regimes, CuBAS constructs compact yet maximally informative training subsets. Extensive evaluation across 60+ benchmark datasets demonstrates consistent and statistically significant improvements over random sampling and uncertainty-based baselines.

Key Technical Details

  • Potts MRF: Isotropic q-state Potts model on k-NN graph; local conditional distribution parameterized by inverse temperature β
  • MPL Estimation: Maximum pseudo-likelihood with secant method for β
  • Curvature Score: S_i(β) = -Ψ_i(β) / (Φ_i(β) + λ), where Φ_i is 1st-order and Ψ_i is 2nd-order observed Fisher information
  • Adaptive Threshold: Otsu's criterion → Ashman's D bimodality check → Tukey-fence fallback
  • Complexity: O(nm log n + nk + T_β nc), linear in k-NN graph edges after construction
  • Results: Mean +13.5pp over entropy sampling at 10% budget across 42 datasets; best on all 60+ datasets