Document Type : Research Article
Author
Department of Statistics, Shahid Beheshti University, Tehran, Iran.
Abstract
In this paper, the claim counts and claim amounts, both with and without missing values, are assumed to be correlated in the context of non-life insurance. The proposed methodology involves fitting joint random effects partially specified models based on the factorization of the joint distribution of claim counts and average claim amounts. Furthermore, the paper introduces joint partially aggregate claims models that account for correlated claim count and amount responses, addressing cases with missing values in both variables and incorporating nonignorable missing data mechanisms. To capture non-linear time effects on the joint model, various spline functions are employed. A full likelihood based estimation procedure is adopted to obtain maximum likelihood estimates for the parameters of the joint partial models involving correlated claim frequencies and severities. For data scenarios where both variables contain nonignorable missing values, a corresponding pure premium formula is derived. The performance of the proposed model is compared to that of the independent aggregate claims model through a simulation study, focusing on differences in the resulting pure premiums. To assess model sensitivity, influence analysis is conducted by examining how small perturbations in the average percent difference (APD) affect likelihood displacement using influence graphs. Finally, the methodology is applied to a real Iranian automobile insurance dataset.
Keywords
[2] E. C. K. Cheung, W. Ni, R. Oh, and J. K. Woo, Bayesian credibility under a bivariate prior on the frequency and the severity of claims Insurance: Mathematics and Economics, (2021) 100, 274-295.
[3] D. R. Cox, and N. Wermuth, Response models for mixed binary and quantitative variables, Biometrika, (1992) 79(3): 441-461.
[4] C. Czado, R. Kastenmeier, E. C. Brechmann, and A. Min, A mixed copula model for insurance claims and claim sizes, Scand. Actuar. J., (2021) 4, 278-305.
[5] C. De Boor, A practical guide to splines, Springer-Verlag, (1987) New York.
[6] R. Derrig, and L. Francis, Distinguishing the Forest from the TREES: A Comparison of Tree Based Data Mining Methods, Casualty Actuarial Society Forum, (2006) 1-49.
[7] R. F. Engle, C. W. J. Granger, J. Rice, and A. Weiss, Semiparametric estimates of the relation between weather and electricity sales, Journal of the American statistical Association, (1986) 81(394), 310–320.
[8] P. H. Eilers, and B. D. Marx, Flexible smoothing with B-splines and penalties, Stat Sci, (1996) 11(2), 89–102.
[9] L. Fahrmeir, and G. Tutz, Multivariate statistical modelling based on generalized linear models, Springer, New york (2001).
[10] R. Fletcher, Practical Methods of Optimization, John Wiley & Sons, New York (2000).
[11] E. W. Frees, J. Gao, and M. A. Rosenberg, Predicting the frequency and amount of health care expenditures, N. Am. Actuar. J. (2011) 15 (3), 377-392.
[12] E. W. Frees, G. Lee, and L. Yang, Multivariate frequency-severity regression models in insurance Risks, (2016) 4(1), 1-36.
[13] J. Garrido, C. Genest, and J. Schulz, Generalized linear models for dependent frequency
and severity of insurance claims, Insurance: Mathematics and Economic, (2016) 70, 205-215.
[14] S. Gschlößl, and C. Czado, Spatial modelling of claim frequency and claim size in non life insurance, Scandinavian Actuarial Journal, ( 2007)3, 202-225.
[15] L. Guelman, and M. Guillén, A causal inference approach to measure price elasticity in Automobile Insurance, Expert Systems with Applications, (2014) 41, 387-396.
[16] L. Goodman, The Variance of the Product of K Random Variables, Journal of the American Statistical Association, (1962) 57(297): 54-60.
[17] D. A. Harvile, and R. W. Mee, A mixed model procedure for analyzing ordered categorical data, Biometrics, (1984) 40, 393-408.
[18] T. J. Hastie, and R. J. Tibshirani, Generalized additive models, CRC press, (1990) 43.
[19] J. J. D. Heckman, Dummy Endogenous variable in a simultaneous Equations system, Econometrica, (1978) 46 (6), 931-59.
compound risk models, SSRN Electronic Journal (2019), https://dx.doi.org/10.2139/ssrn.3045360.
[21] B. Jϕrgensen, M. C. P. de Souza, Fitting Tweedie’s compound Poisson model to insurance claims data, Scand. Actuar. J., (1994) 1, 69-93.
[22] H. Joe, Approximations multivariate normal rectangle probabilities based on conditional expectation, Journal of the American Statistical Association, (1995) 90: 957-967.
[23] R. J. Little, and M. Schluchter, Maximum likelihood estimation for mixed continuous and categorical data with missing values, Biometrika, (1987) 72, 497-512.
[24] R. Oh, P. Shi, and J. Ahn, Bonus-Malus premiums under the dependent frequency severity modeling, Scandinavian Actuarial Journal, (2020) 3, 172-195.
[25] O. A. Quijano-Xacur, and J. Garrido, Generalised linear models for aggregate claims: To Tweedie or not?, Eur. Actuar. J., (2015) 5 (1), 181-202.
[26] A. E. Renshaw, Modelling the claims process in the presence of covariates, ASTIN Bull, (1994) 24 (2), 265-285.
[27] D. B. Rubin, Inference and missing data, Biometrica, (1976) 82, 669-710.
[28] G. Shmueli, T. P. Minka, and J. B. Kadane, S. Borle, and P. Boatwright, A useful distribution for fitting discrete data: Revival of the Conway-Maxwell-Poisson distribution, Appl. Stat. (2005) 54, 127-142.
[29] K. F. Sellers, S. Borle, and G. Shmueli, (2012). The COM-Poisson model for count data: A survey of methods and applications, Appl. Stoch. Model. Bus. (2012) 28, 104-116.
[30] J. C. Stone, Additive regression and other nonparametric models, The Annals of Statistics. (1985) 13(2). 689–705.
[31] G. Tutz, Regression for Categorical Data, Cambridge University Press, Cambridge (2012).
[32] G. Verbeke, and G. Molenberghs, Linear mixed models in practice: A SAS Oriented Approach, Springer (1997).
[33] N. Wang, L. Qian, N. Zhang, and Z. Liu, Modelling the aggregate loss for insurance claims with dependence, Communications in Statistics - Theory and Methods, (2021) 50(9), https://doi.org/10.1080/03610926.2019.1659368.
[34] T. H. Boukadoum, and K. Boukhetala, A Stochastic Process Perspective on Hybrid Log-Normal and Machine Learning Models for Financial Risk under Left-Censored Data, Journal of Mathematics and Modeling in Finance, Allameh Tabataba’i University Press, (2026) 6(1).