TY - GEN
T1 - Predicting News Click-Through Rates Using Transformer-Based Multimodal Models
AU - Liao, Kuoan
AU - Chou, Tzren Ru
N1 - Publisher Copyright:
© 2025 Copyright held by the owner/author(s).
PY - 2026/3/16
Y1 - 2026/3/16
N2 - In the digital era driven by social media, the click-through rate (CTR) of news articles is a crucial indicator of audience engagement, platform traffic, advertisement revenue, and recommendation system ranking. Previous studies primarily focused on predicting CTR before or after publication, typically relying on finished content or early click data, limiting the model’s influence on content creation. This study introduces a “pre-writing prediction” approach, forecasting a news article’s potential CTR even before final titles or content are confirmed. By leveraging initial semantic cues from draft titles, image materials, and scheduled publishing time, the model offers real-time feedback during the content creation process. This mechanism enhances AI-assisted writing and editorial decision-making. This work proposes a multimodal regression model that integrates semantic embeddings (e.g., BERT) and visual encoders (e.g., CLIP) to capture textual, visual, and temporal features. Our experiments evaluate the impact of various input features and model configurations on CTR prediction accuracy. The proposed approach serves as a tool for content generation optimization, media strategy planning, and adaptive AI learning.
AB - In the digital era driven by social media, the click-through rate (CTR) of news articles is a crucial indicator of audience engagement, platform traffic, advertisement revenue, and recommendation system ranking. Previous studies primarily focused on predicting CTR before or after publication, typically relying on finished content or early click data, limiting the model’s influence on content creation. This study introduces a “pre-writing prediction” approach, forecasting a news article’s potential CTR even before final titles or content are confirmed. By leveraging initial semantic cues from draft titles, image materials, and scheduled publishing time, the model offers real-time feedback during the content creation process. This mechanism enhances AI-assisted writing and editorial decision-making. This work proposes a multimodal regression model that integrates semantic embeddings (e.g., BERT) and visual encoders (e.g., CLIP) to capture textual, visual, and temporal features. Our experiments evaluate the impact of various input features and model configurations on CTR prediction accuracy. The proposed approach serves as a tool for content generation optimization, media strategy planning, and adaptive AI learning.
KW - BERT
KW - click-through rate prediction
KW - CLIP
KW - news headline
KW - Transformer
UR - https://www.scopus.com/pages/publications/105035380650
UR - https://www.scopus.com/pages/publications/105035380650#tab=citedBy
U2 - 10.1145/3772673.3772708
DO - 10.1145/3772673.3772708
M3 - Conference contribution
AN - SCOPUS:105035380650
T3 - ACMLC 2025 - Proceedings of 2025 7th Asia Conference on Machine Learning and Computing
SP - 28
EP - 33
BT - ACMLC 2025 - Proceedings of 2025 7th Asia Conference on Machine Learning and Computing
PB - Association for Computing Machinery, Inc
T2 - 2025 7th Asia Conference on Machine Learning and Computing, ACMLC 2025
Y2 - 25 July 2025 through 27 July 2025
ER -