УКР ENG

Search:


Email:  
Password:  

 REGISTRATION CERTIFICATE

KV #19905-9705 PR dated 02.04.2013.

 FOUNDERS

RESEARCH CENTER FOR INDUSTRIAL DEVELOPMENT PROBLEMS of NAS of Ukraine (KHARKIV, UKRAINE)

ROR

EDRPOU 05481984

According to the decision No. 802 of the National Council of Television and Radio Broadcasting of Ukraine dated 14.03.2024, is registered as a subject in the field of print media.
ID R30-03156

 PUBLISHER

Liburkina L. M.

 CATALOG

Annotated catalogue (2011)
Annotated catalogue (2012)
Annotated catalogue (2013)
Annotated catalogue (2014)
Annotated catalogue (2015)
Annotated catalogue (2016)
Annotated catalogue (2017)
Annotated catalogue (2018)
Annotated catalogue (2019)
Annotated catalogue (2020)
Annotated catalogue (2021)
Annotated catalogue (2022)
Annotated catalogue (2023)
Annotated catalogue (2024)
Annotated catalogue (2025)
Annotated catalogue (2026)
Thematic sections of the journal
Proceedings of scientific conferences


Impact of Structured Outliers on Regression Inference
Manzhos T. V., Melnyk O. O.

Manzhos, Tetіana V., and Melnyk, Olha O. (2026) “Impact of Structured Outliers on Regression Inference.” Business Inform 6:233–244.
https://doi.org/10.32983/2222-4459-2026-6-233-244

Section: Economic and Mathematical Modeling

Article is written in English
Downloads/views: 0

Download article (pdf) -

UDC 330.43:519.237.5

Abstract:
The article explores the problem of the reliability of regression analysis results in the presence of atypical observations. In econometric studies, such observations can occur in various types of data: macroeconomic, financial, regional, sectoral, market, or broader socioeconomic data. Their appearance is not always related to random errors or technical inaccuracies. Often, they reflect crisis periods, structural shifts, the heterogeneity of the studied objects, the concentration of certain indicators, or the presence of local patterns in the data. Under these conditions, it is important not only to detect outliers but also to understand how their nature affects coefficient estimation, statistical conclusions, and the economic interpretation of the model. The article aims to compare the consequences of different types of outliers for estimating the parameters of a linear regression. The article considers five scenarios: random vertical outliers, non-random vertical outliers, high-leverage observations that align with the basic regression relationship, high-leverage observations that distort this relationship, and clustered high-leverage observations that might form a separate local regime. The study is based on Monte Carlo simulations with 1000 repetitions. The base model includes three regressors with a given correlation structure; one of the coefficients is zero. This allows for assessing not only the bias of the main coefficient but also the appearance of a false effect for a variable that, according to the model, does not affect the outcome. The article compares three evaluation approaches: the usual least squares method, Huber regression, and the least trimmed squares method. The quality of evaluation is analyzed based on coefficient bias, root mean square error, sign stability, the size of false effects, and the frequency of large false effects. Simulation results show that the impact of outliers on a regression model is determined not only by their presence but also by how they actually occur. Vertical outliers primarily increase the variability of residuals and are relatively well handled by Huber regression. On the other hand, high-leverage observations require separate interpretation: if they follow the underlying regression relationship, they should not be automatically considered problematic. But if such observations change the slope of that relationship, they can significantly bias coefficient estimates and create false effects for correlated regressors. In such cases, the least trimmed squares method turns out to be more effective than Huber regression. At the same time, clustered observations with high leverage remain difficult to correct even after trimming part of the data, which may indicate structural heterogeneity in the sample. The scientific novelty of the article lies in the fact that different types of atypical observations are compared not only by their presence but also by the nature of their impact on econometric estimation. The practical value of the results lies in forming guidelines for choosing a robust method. If the problem is mainly related to large residuals, Huber regression is appropriate. If the atypical observations have high leverage and alter the shape of the regression relationship, it is more reasonable to use trimming methods or conduct additional subgroup and local regime analyses.

Keywords: regression inference; structured outliers; high-leverage observations; least squares method; Huber regression; least trimmed squares method; Monte Carlo modeling; econometric estimation.

Fig.: 4. Tabl.: 3. Formulae: 17. Bibl.: 16.

Manzhos Tetіana V. – Candidate of Sciences (Physics and Mathematics), Associate Professor, Associate Professor, Department of Higher Mathematics, Kyiv National Economic University named after Vadym Hetman (54/1 Beresteiskyi Ave., Kyiv, 03057, Ukraine)
Email: [email protected]
Melnyk Olha O. – Candidate of Sciences (Physics and Mathematics), Associate Professor, Associate Professor, Department of Higher Mathematics, Kyiv National Economic University named after Vadym Hetman (54/1 Beresteiskyi Ave., Kyiv, 03057, Ukraine)
Email: [email protected]

List of references in article

Atkinson A. C. & Riani M. (2000). Robust Diagnostic Regression Analysis. New York: Springer. https://doi.org/10.1007/978-1-4612-1160-0
Belsley D. A., Kuh E. & Welsch R. E. (1980). Regression Diagnostics: Identifying Influential Data and Sources of Collinearity. New York: Wiley. https://doi.org/10.1002/0471725153
Cook R. D. (1977). Detection of influential observation in linear regression. Technometrics, 1(19), 15–18. https://doi.org/10.1080/00401706.1977.10489493
Gao Z. & Moon H. R. (2024). Robust estimation of regression models with potentially endogenous outliers via a modern optimization lens. arXiv preprint. https://doi.org/10.48550/arXiv.2408.03930
Hampel F. R., Ronchetti E. M., Rousseeuw P. J. & Stahel W. A. (1986). Robust Statistics: The Approach Based on Influence Functions. New York: Wiley. https://doi.org/10.1002/9781118186435
Huber P. J. (1973). Robust regression: asymptotics, conjectures and Monte Carlo. The Annals of Statistics, 5(1), 799–821. https://doi.org/10.1214/aos/1176342503
Koenker R. & Bassett G. (1978). Regression quantiles. Econometrica, 1(46), 33–50. https://doi.org/10.2307/1913643
Leone A. J., Minutti-Meza M. & Wasley C. E. (2019). Influential observations and inference in accounting research. The Accounting Review, 6(94), 337–364. https://doi.org/10.2308/accr-52396
Maronna R. A., Martin R. D. & Yohai V. J. (2006). Robust Statistics: Theory and Methods. Chichester: Wiley. https://doi.org/10.1002/0470010940
Rousseeuw P. J. & Hubert M. (2018). Anomaly detection by robust statistics. WIREs Data Mining and Knowledge Discovery, 2(8), e1236. https://doi.org/10.1002/widm.1236
Rousseeuw P. J. (1984). Least median of squares regression. Journal of the American Statistical Association, 388(79), 871–880. https://doi.org/10.1080/01621459.1984.10477105
Rousseeuw P. J. & Leroy A. M. (1987). Robust Regression and Outlier Detection. New York: Wiley. https://doi.org/10.1002/0471725382
Rousseeuw P. J. & van Zomeren B. C. (1990). Unmasking multivariate outliers and leverage points. Journal of the American Statistical Association, 411(85), 633–639. https://doi.org/10.1080/01621459.1990.10474920
Rousseeuw P. J. & Van Driessen K. (2006). Computing LTS regression for large data sets. Data Mining and Knowledge Discovery, 12, 29–45. https://doi.org/10.1007/s10618-005-0024-4
Yohai V. J. (1987). High breakdown-point and high efficiency robust estimates for regression. The Annals of Statistics, 2(15), 642–656. https://doi.org/10.1214/aos/1176350366
Zaman A., Rousseeuw P. J. & Orhan M. (2001). Econometric applications of high-breakdown robust regression techniques. Economics Letters, 1(71), 1–8. https://doi.org/10.1016/S0165-1765(00)00404-3

 FOR AUTHORS

License Contract

Conditions of Publication

Article Requirements

Regulations on Peer-Reviewing

Current Issue

Frequently asked questions

 INFORMATION

Main page

Editorial staff

Editorial policy

About the Journal

Aim and Scope

AI and generative AI tools policy

Announcements and news

The Plan of Scientific Conferences

Indexing

 OUR PARTNERS

Journal «The Problems of Economy»

  © Business Inform, 1992 - 2026 The site and its metadata are licensed under CC BY-SA. Write to webmaster