papersTODAY 04:00 UTC
Diffusion Model Imputes Missing Values in Mixed Numerical and Categorical Data
A new arXiv paper introduces Impute-EM, a diffusion-based approach for filling in missing values in datasets that mix numerical, categorical, and binary variables. Most existing diffusion imputation methods convert discrete variables into continuous stand-ins, which the authors argue is a limitation the native mixed-state design avoids. The work targets heterogeneous data mining settings where such mixed variable types commonly appear together.