Go top
Conference paper information

SemTab: A Hybrid Framework for Semantic Feature Generation on Tabular Data

O. Chen, K. Chou, R. Nagpal, R. Palacios, A. Gupta

Undergraduate Research Technology Conference - MIT URTC 2025, Cambridge (United States of America). 10-12 October 2025


Summary:

Machine learning models on tabular datasets often struggle to understand the context between features, which can limit their accuracy. We propose SemTab, a hybrid framework for generating semantic features that utilizes an open-source Large Language Model (LLM). We evaluated our framework using three benchmark datasets: Adult Income, German Credit, and Bank Marketing. We compared its performance against several off-the-shelf LLMs. The results show that SemTab achieved the highest accuracy across all the classification tasks. For instance, on the Bank Marketing dataset, SemTab achieved an accuracy of 8 0%, which is approximately 2 0% improvement over the baseline models. This work highlights that a hybrid architecture is a practical approach for applying language models to structured tabular data, yielding accurate and interpretable results for various downstream tasks.


Keywords: Tabular Data, Semantic Feature Generation, LLMs, Model Interpretability


DOI: DOI icon https://doi.org/10.1109/URTC68753.2025.11533131

Published in: 2025 IEEE MIT Undergraduate Research Technology Conference (URTC), pp: 1-5, ISBN: 979-8-3315-5938-0

Publication date: 26-May-2026.


Citation:
O. Chen, K. Chou, R. Nagpal, R. Palacios, A. Gupta, "SemTab: A Hybrid Framework for Semantic Feature Generation on Tabular Data", presented at Undergraduate Research Technology Conference - MIT URTC 2025, Cambridge, United States of America, 10-12 October 2025. In: 2025 IEEE MIT Undergraduate Research Technology Conference (URTC), pp. 1-5, doi: 10.1109/URTC68753.2025.11533131

    Research groups:
  • Instituto de Investigación Tecnológica (IIT)

IIT-25-413C

Request Request the document to be emailed to you.