Undergraduate Research Technology Conference - MIT URTC 2025, Cambridge (United States of America). 10-12 October 2025
Summary:
Machine learning models on tabular datasets often struggle to understand the context between features, which can limit their accuracy. We propose SemTab, a hybrid framework for generating semantic features that utilizes an open-source Large Language Model (LLM). We evaluated our framework using three benchmark datasets: Adult Income, German Credit, and Bank Marketing. We compared its performance against several off-the-shelf LLMs. The results show that SemTab achieved the highest accuracy across all the classification tasks. For instance, on the Bank Marketing dataset, SemTab achieved an accuracy of 8 0%, which is approximately 2 0% improvement over the baseline models. This work highlights that a hybrid architecture is a practical approach for applying language models to structured tabular data, yielding accurate and interpretable results for various downstream tasks.
Keywords: Tabular Data, Semantic Feature Generation, LLMs, Model Interpretability
DOI:
https://doi.org/10.1109/URTC68753.2025.11533131
Published in: 2025 IEEE MIT Undergraduate Research Technology Conference (URTC), pp: 1-5, ISBN: 979-8-3315-5938-0
Publication date: 26-May-2026.
Citation:
O. Chen, K. Chou, R. Nagpal, R. Palacios, A. Gupta, "SemTab: A Hybrid Framework for Semantic Feature Generation on Tabular Data", presented at Undergraduate Research Technology Conference - MIT URTC 2025, Cambridge, United States of America, 10-12 October 2025. In: 2025 IEEE MIT Undergraduate Research Technology Conference (URTC), pp. 1-5, doi: 10.1109/URTC68753.2025.11533131
IIT-25-413C