ConceptioArchiveNCBI PubMed Central
NCBI PubMed Centralopen access

Regression for Left-Truncated and Right-Censored Data: A Semiparametric Sieve Likelihood Approach.

Matthews S et al. · ncbi_pmc
NCBI PubMed Central · Papers · License: Open Access
Open Source ↗Direct PDF ↓
computer-science-education
computer science education

Regression for Left‐Truncated and Right‐Censored Data: A Semiparametric Sieve Likelihood Approach - PMC Skip to main content An official website of the United States government Here's how you know Here's how you know Official websites use .gov A .gov website belongs to an official government organization in the United States. Secure .gov websites use HTTPS A lock ( Lock Locked padlock icon ) or https:// means you've safely connected to the .gov website. Share sensitive information only on official, secure websites. Search Log in Dashboard Publications Account settings Log out Search… Search NCBI Primary site navigation Search Logged in as: Dashboard Publications Account settings Log in Search PMC Full-Text Archive Search in PMC Journal List User Guide PERMALINK Copy As a library, NLM provides access to scientific literature. Inclusion in an NLM database does not imply endorsement of, or agreement with, the contents by NLM or the National Institutes of Health. Learn more: PMC Disclaimer | PMC Copyright Notice Stat Med . 2026 Apr 2;45:e70509. doi: 10.1002/sim.70509 Search in PMC Search in PubMed View in NLM Catalog Add to search Regression for Left‐Truncated and Right‐Censored Data: A Semiparametric Sieve Likelihood Approach Spencer Matthews Spencer Matthews 1 Department of Statistics, University of California, Irvine, California, USA Find articles by Spencer Matthews 1 , Bin Nan Bin Nan 1 Department of Statistics, University of California, Irvine, California, USA Find articles by Bin Nan 1, ✉ Author information Article notes Copyright and License information 1 Department of Statistics, University of California, Irvine, California, USA * Correspondence: Bin Nan ( [email protected] ) ✉ Corresponding author. Revised 2026 Jan 20; Received 2025 Feb 12; Accepted 2026 Mar 16; Issue date 2026 Apr. © 2026 The Author(s). Statistics in Medicine published by John Wiley & Sons Ltd. This is an open access article under the terms of the http://creativecommons.org/licenses/by-nc/4.0/ License, which permits use, distribution and reproduction in any medium, provided the original work is properly cited and is not used for commercial purposes. PMC Copyright notice PMCID: PMC13048288  PMID: 41923496 ABSTRACT Cohort studies of the onset of a disease often encounter left‐truncation on the event time of interest in addition to right‐censoring due to variable enrollment times of study participants. Analysis of such event time data can be biased if left‐truncation is not handled properly. We propose a semiparametric sieve likelihood approach for fitting a linear regression model to data where the response variable is subject to both left‐truncation and right‐censoring. We show that the estimators of regression coefficients are consistent, asymptotically normal and semiparametrically efficient. Extensive simulation studies show the effectiveness of the method across a wide variety of error distributions. We further illustrate the method by analyzing datasets from the Canadian Study of Health and Aging and The 90+ Study for aging and dementia. Keywords: accelerated failure time model, Alzheimer's disease, B‐spline, bundled parameters, dementia, semiparametric efficiency 1. Introduction There has been significant research effort in understanding age‐related diseases, such as Alzheimer's disease and dementia, in recent decades. A key method in understanding these diseases is to design a prospective cohort study, where study participants are recruited and followed, with data collected at regular intervals. Oftentimes the ultimate outcome of interest is the time at which the disease occurs. Such studies typically require that participants who enroll in the study meet an age requirement while being disease free; for instance, The 90+ Study only enrolled dementia‐free individuals at age 90 or older [ 1 ]. This method of participant enrollment induces left‐truncation into the resulting dataset in addition to right‐censoring. Left‐truncation, often called delayed entry in longitudinal studies, occurs when individuals who have already developed the condition of interest prior to the commencement of the study are not observed or are excluded from the analysis. Simply ignoring left‐truncation introduces bias, as individuals who develop the disease earlier are systematically under represented, potentially skewing the findings toward a healthier or longer‐lived population. Appropriately accounting for left‐truncation is therefore essential to obtain valid inference, particularly in longitudinal studies of aging and chronic disease. Another common feature of event time data is right‐censoring, which occurs when some participants in the study do not experience the event of interest within the specified study time period, or are lost to follow‐up before the event can be recorded. As a result, the exact time of the event for such an individual remains unknown, and the available information is limited to the fact that the event has not occurred up to the last observed time. While fitting parametric models or semiparametric hazard regression models, for example, the Cox model, to data subject to left‐truncation and right‐censoring has been studied extensively [ 2 ], the case of fitting a semiparametric linear model to such data is underdeveloped, and a linear model is a natural choice for the transformed outcome of interest in the example of The 90+ Study. For some reference time A 0 , often referred to as time zero in survival analysis which can be either deterministic or random, define ( Y ˜ , T ˜ , C ˜ ) as the underlying event time, left‐truncation time (i.e., study entry time), and right‐censoring time since A 0 , respectively, where Y ˜ can be left‐truncated by T ˜ and right‐censored by C ˜ . Given a well‐defined continuous and monotonic transformation ξ ( · ) , let Y ∗ = ξ ( Y ˜ ) , T = ξ ( T ˜ ) and C = ξ ( C ˜ ) , then Y ∗ can be left‐truncated by T and right‐censored by C . Define Y = Y ∗ ∧ C , where a ∧ b = min ( a , b ) , and Δ = 1 ( Y ∗ ≤ C ) , where 1 ( · ) is an indicator function. We also define a ∨ b = max ( a , b ) . Let X denote a set of covariates. Variables ( Y , Δ , X , T ) become observed when Y > T and otherwise unobservable, giving the so‐called left‐truncated data that are also subject to right‐censoring, which follow a conditional distribution given Y > T . For the underlying transformed event time of interest Y ∗ , consider the following linear regression model given covariates X : Y ∗ = X T β + e , (1) where β ∈ R d , d is a fixed integer, and the error e follows some unknown distribution F e with density f e . Thus model ( 1 ) is a semiparametric model. If ξ ( · ) = log ( · ) , then this becomes the so‐called accelerated failure time model as in Kalbfleisch and Prentice [ 3 ], chapter 7. We extend the terminology to model ( 1 ) with any known transformation ξ . Another model of potential interest similar to the above model ( 1 ) is the linear transformation model, where the function ξ is unknown but the distribution of e is given, see Wang and Wang [ 4 ] and references of earlier work therein. Incorporating left‐truncation in this setting would be an interesting future work. We consider a known transformation ξ here for a more straightforward interpretation of β . One method for fitting a semiparametric model like ( 1 ) is rank regression. Adichie [ 5 ] was among the first who used residual rank statistics for fitting regression parameters, and Bhattacharya et al. [ 6 ] extended the approach to left‐truncated data in their work in astronomy. Lai and Ying [ 7 , 8 , 9 ] used these ideas in exploring rank statistics with left‐truncated and right‐censored data as they studied the behavior and asymptotic properties of such a regression. Kim and Lai [ 10 ] pointed out various issues with practical implementations of rank regression methods, the most important being the large computational cost associated with fitting these models. In an attempt to overcome some of these issues, they introduced an adaptive M‐estimation method using efficient scores for censored and truncated regression models in which they used splines to approximate their estimating equations. While they showed their method to have nice theoretical properties, the implementation is dense and no code package is readily available. Also note that another estimation method for model ( 1 ) under left‐truncation and right‐censoring is the length‐biased sample approach, where it is assumed that the left‐truncation time follows a uniform distribution [ 11 , 12 , 13 ]. Following Wang [ 14 ], in contrast, the method we consider in this work does not make any distributional assumption about the left‐truncation time. A more computationally efficient approach, and statistically efficient as well, is to fit model ( 1 ) using a sieve likelihood. The method of sieves provides a general framework for estimating infinite‐dimensional parameters by approximating them with a sequence of finite‐dimensional spaces whose complexity increases with the sample size. First explored by Geman and Hwang [ 15 ], sieve estimation has developed into an effective tool for semiparametric and nonparametric inference. It allows the flexibility of infinite‐dimensional models while preserving the tractability of likelihood‐based estimation. In the sieve likelihood approach, the likelihood function is maximized over a sequence of finite‐dimensional approximations of infinite‐dimensional parameters with dimension growing at an appropriate rate to ensure both consistency and asymptotic efficiency [ 16 , 17 ]. This framework has been successfully applied in a variety of semiparametric contexts, including censored and truncated data models [ 18 ], where traditional likelihood methods can be difficult to implement due to the presence of infinite‐dimensional nuisance parameters. In this work, we consider a sieve likelihood approach for fitting model ( 1 ) to left‐truncated and right‐censored data where the unknown infinite‐dimensional nuisance parameter, the logarithm of the hazard function of e in our case, is estimated using splines. Our method builds on the work of Ding and Nan [ 19 ] for bundled parameters where only right‐censoring was considered. This approach allows for a semiparametric likelihood estimation for a frequently encountered problem in practice, which not only provides a simple numerical implementation but also yields desirable asymptotic properties for the regression coefficient estimates, namely asymptotical normality and semiparametric efficiency. We demonstrate this approach through simulation studies and applications to the Canadian Study of Health and Aging [ 20 , 21 ] and The 90+ Study [ 1 ]. 2. Sieve Likelihood Approach A sieve approach approximates the true unknown infinite‐dimensional parameter with a sequence of increasingly complex parametric models. In our case, we propose to use a sequence of B‐spline basis expansions to approximate the logarithm of the unknown hazard function of e in model ( 1 ). Denote the hazard function of e by λ e , and let g = log λ e . For a finite interval [ a , b ] and K knots placed within the interval: a < κ 1 < ⋯ < κ K < b , we approximate g by g ( s ) = log λ e ( s ) ≈ ∑ k = 1 K γ k B k ( s ) , where B k , k = 1 , … , K , are cubic B‐spline basis functions. We chose cubic B‐splines so that the second derivative is continuous, but other splines may be considered as well. For further discussion see De Boor [ 22 ], chapter 9. Since e = Y ∗ − X T β , g is a composite function of β . Thus ( β , g ) form a pair of bundled parameters [ 19 ]. Denote ϵ β , i = Y i − X i T β and τ β , i = T i − X i T β that denotes the left‐truncation time on the residual scale. Assume Y ∗ is independent of ( C , T ) given X . Then for independent and identically distributed observations { ( y i , δ i , x i , t i ) : y i > t i , i = 1 , … , n } , we have the following likelihood function: ∏ i = 1 n exp { g ϵ β , i } Δ i exp − ∫ τ β , i ϵ β , i exp { g ( u ) } d u , which is then approximated by the sieve likelihood: L n ( β , γ ) = ∏ i = 1 n exp ∑ k = 1 K γ k B k ϵ β , i Δ i × exp − ∫ τ β , i ϵ β , i exp ∑ k = 1 K γ k B k ( u ) d u . See the Appendix for the detailed construction of the density function of the observed data ( Y , Δ , X , T ) given Y > T which is obtained by using the left‐truncation observation condition and leads to the above likelihood function for data with left‐truncation and right‐censoring. Then we obtain the following sieve log‐likelihood function: ℓ n ( β , γ ) = ∑ i = 1 n Δ i ∑ k = 1 K γ k B k ϵ β , i − ∫ τ β , i ϵ β , i exp ∑ k = 1 K γ k B k ( s ) d s . (2) Denote B ˙ k ( s ) = d B k ( s ) / d s and B ¨ k = d 2 B k ( s ) / d s 2 . Taking partial derivatives of the above sieve log‐likelihood function ( 2 ), we arrive at the following sieve score functions: S n ( β , γ ) = ∂ ℓ n ( β , γ ) ∂ β , ∂ ℓ n ( β , γ ) ∂ γ T , (3) with ∂ ℓ n ( β , γ ) ∂ β = ∑ i = 1 n X i T − Δ i ∑ k = 1 K γ k B ˙ k ϵ β , i + exp ∑ k = 1 K γ k B k ϵ β , i − exp ∑ k = 1 K γ k B k τ β , i (4) and ∂ ℓ n ( β , γ ) ∂ γ k = ∑ i = 1 n Δ i B k ϵ β , i − ∫ τ β , i ϵ β , i B k ( s ) exp ∑ k = 1 K γ k B k ( s ) d s . (5) Furthermore, the observed sieve information matrix is I n ( β , γ ) = ∂ 2 ℓ n ( β , γ ) ∂ β ∂ β T ∂ 2 ℓ n ( β , γ ) ∂ β ∂ γ T ∂ 2 ℓ n ( β , γ ) ∂ γ ∂ β T ∂ 2 ℓ n ( β , γ ) ∂ γ ∂ γ T , (6) with ∂ 2 ℓ n ( β , γ ) ∂ β ∂ β T = ∑ i = 1 n X i X i T Δ i ∑ k = 1 K γ k B ¨ k ϵ β , i − exp ∑ k = 1 K γ k B k ϵ β , i ∑ k = 1 K γ k B ˙ k ( ϵ β , i ) + exp ∑ k = 1 K γ k B k τ β , i ∑ k = 1 K γ k B ˙ k τ β , i , ∂ 2 ℓ n ( β , γ ) ∂ β ∂ γ k = ∑ i = 1 n X i T − Δ i B ˙ k ϵ β , i + B k ϵ β , i exp ∑ k = 1 K γ k B k ϵ β , i − B k τ β , i exp ∑ k = 1 K γ k B k τ β , i , ∂ 2 ℓ n ( β , γ ) ∂ γ k ∂ γ l = ∑ i = 1 n − ∫ τ β , i ϵ β , i B k ( s ) B l ( s ) exp ∑ m = 1 K γ m B m ( s ) d s . We consider a gradient search algorithm by setting ( 4 ) and ( 5 ) equal to zero and solving for β and γ , which leads to the sieve maximum likelihood estimates ( β ^ n , γ ^ n ) . Similar to Ding and Nan [ 19 ] we can see that, under the regularity conditions given in the next section, the sieve log‐likelihood function ( 2 ) is locally concave in a neighborhood of true parameters ( β 0 , g 0 ) . Thus, we start the search algorithm from multiple randomly selected initial values and set ( β ^ n , γ ^ n ) to be the converging result that gives the maximum value of the sieve log‐likelihood ( 2 ). Specifically in our numerical examples, we generate 10 random initial values each obtained by adding a vector of random noises to the naive ordinary least squares estimator of β obtained from complete observations, where the standard deviation of each random noise is chosen to be three times the standard error of the corresponding estimated coefficient from the naive model. We multiply the standard error by three to avoid getting caught too close to the naive estimate for all random starts. We sample initial γ values from the standard normal distribution. To avoid singular information matrix during search iterations, we only use diagonal elements of the sieve information matrix ( 6 ), with the final full information matrix evaluated at ( β ^ n , γ ^ n ) . The method works well in our simulation studies presented in Section 4 . 3. Asymptotic Results Define ϵ 0 = Y − X T β 0 and τ 0 = T − X T β 0 , where β 0 is the true parameter. Let λ 0 be the hazard function of e 0 = Y ∗ − X T β 0 and Λ 0 be the corresponding cumulative hazard function. Following Lai and Ying [ 7 ] we assume e 0 is independent of ( C , T , X ) , a reasonable assumption for a linear regression model, which implies that Y ∗ is independent of ( C , T ) given X , an assumption of conditional independent left‐truncation and right‐censoring that we need for obtaining the joint density function of the observed data ( Y , Δ , X , T ) with Y > T in the Appendix. To obtain desirable asymptotic properties of the proposed estimator, we introduce the same regularity conditions in Ding and Nan [ 19 ] with a few additional assumptions for the left‐truncation. A.1 The true parameter β 0 belongs to the interior of a compact set ℬ ⊆ R d . A.2 (a) The covariate X takes values in a bounded subset 𝒳 ⊆ R d ; (b) E ( X X T ) is non‐singular. A.3 (a) There is a constant b < ∞ such that, for some constant δ , pr ( ϵ 0 > b | X ) ≥ δ > 0 almost surely with respect to the probability measure of X . This implies that Λ 0 ( b ) ≤ − log δ < ∞ ; (b) We also assume pr ( τ 0 < b | X ) ≥ 1 − δ 1 for some constant δ 1 ∈ ( 0 , 1 ) almost surely, and pr ( τ 0 > ϵ 0 | X ) ≤ δ 2 for some constant δ 2 ∈ [ 0 , 1 ) almost surely. A.4 The error e 0 's density f 0 and its derivative f ˙ 0 are bounded and ∫ f ˙ 0 ( s ) / f 0 ( s ) 2 f 0 ( s ) d s < ∞ . A.5 (a) The conditional density of C given ( T , X ) and its derivative are uniformly bounded for all possible values of T ∈ 𝒯 ⊆ R and X ∈ 𝒳 , that is, sup t ∈ 𝒯 , x ∈ 𝒳 f C | T , X ( y | t , x ) ≤ K 1 , sup t ∈ 𝒯 , x ∈ 𝒳 f ˙ C | T , X ( y | t , x ) ≤ K 2 for all y ∈ 𝒴 ⊆ R with some constants K 1 , K 2 > 0 ; (b) Similarly, the conditional density of T given X and its derivative are uniformly bounded for all possible values of X ∈ 𝒳 , that is, sup x ∈ 𝒳 f T | X ( t | x ) ≤ K 3 , sup x ∈ 𝒳 f ˙ T | X ( t | x ) ≤ K 4 for all t ∈ 𝒯 with some constants K 3 , K 4 > 0 ; (c) Both f C | T , X and f T , X | Y > T are free of parameters β and λ e . A.6 Let 𝒢 p denote the collection of bounded functions g on [ a , b ] with bounded derivatives g ( j ) , j = 1 , … , k . The k th derivative g ( k ) satisfies the following Lipschitz continuity condition: g ( k ) ( s ) − g ( k ) ( t ) ≤ L | s − t | m for s , t ∈ [ a , b ] , where k is a positive integer and m ∈ ( 0 , 1 ] such that p = k + m ≥ 3 , L < ∞ is an unknown constant, a = inf y , x ( y − x T β 0 ) ∨ inf t , x ( t − x T β 0 ) > − ∞ and b is defined in condition (A.3). The true log hazard function of the error e 0 , g 0 ( · ) = log λ 0 ( · ) , belongs to 𝒢 p . A.7 For some η ∈ ( 0 , 1 ) , u T Var ( X | ϵ 0 ) u ≥ η u T E ( X X T | ϵ 0 ) u almost surely for all u ∈ R d . The modified regularity conditions to those given in Section 3 of Ding and Nan [ 19 ] are in (A.3) (b) and (A.5). Condition (A.3) (b) guarantees that we observe left‐truncated data almost surely. Condition (A.5) (a) is a natural extension of the condition in Ding and Nan [ 19 ] for the conditional density function of C . Condition (A.5) (b) plays a similar role as (A.5) (a). Condition (A.5) (c) contains a noninformative censoring assumption for right‐censored data and a similar but less stringent noninformative truncation assumption made by Lai and Ying [ 23 ] for left‐truncated data. Condition (A.5) (c) simplifies the likelihood function to the desirable form we use in this article. Note that while f T , X | Y > T being free of β and λ e holds in our setup (as shown in Wang [ 14 ]), we include this assumption in our conditions for clarity. When sup { t ∈ 𝒯 } ≤ inf { y ∈ 𝒴 } , left‐truncation does not play any role so we only deal with right‐censored data and the problem reduces to that in Ding and Nan [ 19 ]. On the other hand, when pr ( Y ∗ ≤ C ) = 1 , the problem reduces to left‐truncation only. Consider the collection of functions ℋ p = ζ ( · , β ) : ζ ( s , x , β ) = g ( ψ ( s , x , β ) ) , g ∈ 𝒢 p , s ∈ [ a , b ] , x ∈ 𝒳 , β ∈ ℬ } , where ψ ( s , x , β ) = s − x T ( β − β 0 ) and 𝒢 p is defined in Condition (A.6). Here ζ is a composite function of g composed with ψ . Note that ζ ( s , x , β 0 ) = g ( s ) . Then for ζ ( · , β ) ∈ ℋ p we define the following norm: ‖ ζ ( · , β ) ‖ 2 = ∫ 𝒳 ∫ a b g s − x T ( β − β 0 ) 2 d Λ 0 ( s ) d F X ( x ) 1 / 2 . For any θ 1 = ( β 1 , ζ 1 ( · , β 1 ) ) and θ 2 = ( β 2 , ζ 2 ( · , β 2 ) ) in the space of Θ p = ℬ × ℋ p , define the following distance: d ( θ 1 , θ 2 ) = | β 1 − β 2 | 2 + ‖ ζ 1 ( · , β 1 ) − ζ 2 ( · , β 2 ) ‖ 2 2 1 / 2 . (7) This is the same distance used by Ding and Nan [ 19 ] in their treatment of the right‐censored case. Let 𝒢 n p be the space of polynomial splines of order p ≥ 1 . For details on this space see Schumaker [ 24 ], chapter 4. Denote ℋ n p = ζ ( · , β ) : ζ ( s , x , β ) = g ( ψ ( s , x , β ) ) , g ∈ 𝒢 n p , s ∈ [ a , b ] , x ∈ 𝒳 , β ∈ ℬ } and Θ n p = ℬ × ℋ n p . Clearly, ℋ n p ⊆ ℋ n + 1 p ⊆ ⋯ ⊆ ℋ p for all n ≥ 1 . The sieve estimator θ ^ n = ( β ^ n , ζ ^ n ( · , β ^ n ) ) , where ζ ^ n ( s , x , β ^ n ) = ĝ n ( s − x T ( β ^ n − β 0 ) ) , is the maximizer of the log‐likelihood over the sieve space Θ n p . The following theorem gives the convergence rate of the proposed estimator θ ^ n to the true parameter θ 0 = ( β 0 , ζ 0 ( · , β 0 ) ) = ( β 0 , g 0 ) . Theorem 1 Let K n be the number of internal knots in the spline space and K n = O ( n ν ) , where ν satisfies 1 / { 2 ( 1 + p ) } < ν < 1 / ( 2 p ) with p being the smoothness parameter defined in condition ( A .6). Suppose conditions ( A .1)–( A .7) hold. Then d ( θ ^ n , θ 0 ) = O p n − min ( p ν , ( 1 − ν ) / 2 ) , where d ( · , · ) is defined in ( 7 ). Remark 1 The sieve space 𝒢 n p does not have to be restricted to the B‐spline space. It can be any sieve space as long as the estimator θ ^ n ∈ ℬ × ℋ n p satisfies the conditions of Theorem 1 in Shen and Wong [ 16 ]. Our choice of the cubic B‐spline space is primarily motivated by its simplicity of numerical implementation, which is a tremendous advantage of the proposed approach over existing numerical methods for fitting accelerated failure time models for data subject to right‐censoring and left‐truncation. The proof of Theorem 1 follows that of Ding and Nan [ 25 ] by verifying the conditions of Shen and Wong [ 16 ]. The main difference is controlling the additional complication brought by τ β = T − X T β in the likelihood function, which is managed by Condition (A.3)(b) as well as desirable properties of linear and indicator functions. Theorem 1 implies that if ν = 1 / ( 1 + 2 p ) , d ( θ ^ n , θ 0 ) = O p ( n − p / ( 1 + 2 p ) ) which is the optimal convergence rate in the nonparametric regression setting. Although the overall convergence rate is slower than n − 1 / 2 , the next theorem states that the proposed estimator of the regression parameter is still asymptotically normal and semiparametrically efficient. Theorem 2 Given the following efficient score function obtained by Lai and Ying [ 23 ] for the censored and truncated linear model : l β 0 ∗ ( Y , Δ , X , T ) = ∫ X − E X | ϵ 0 ≥ s , τ 0 ≤ s − λ ˙ 0 λ 0 ( s ) d M ( s ) , where M ( s ) = Δ 1 ( ϵ 0 ≤ s ) 1 ( ϵ 0 ≥ τ 0 ) − ∫ − ∞ s 1 ( ϵ 0 ≥ u ) 1 ( τ 0 ≤ u ) λ 0 ( u ) d u is the failure counting process martingale, suppose that the conditions in Theorem 1 hold and I ( β 0 ) = P { l β 0 ∗ ( Y , Δ , X , T ) ⊗ 2 } is nonsingular. Then n 1 / 2 ( β ^ n − β 0 ) = n − 1 / 2 I − 1 ( β 0 ) ∑ i = 1 n l β 0 ∗ ( Y i , Δ i , X i , T i ) + o p ( 1 ) → N ( 0 , I − 1 ( β 0 ) ) in distribution . The proof of Theorem 2 closely follows the work of Ding and Nan [ 19 , 25 ] and Zhao et al. [ 26 ] by checking Conditions (A1) through (A6) of Ding and Nan [ 19 ]. Verifying their Condition (A3) requires searching for the least‐favorable submodel, which leads to the same efficient score function found by Lai and Ying [ 23 ]. The main difference to Ding and Nan [ 19 ] is again the inclusion of the left‐truncation indicator, which has minimum impact under the additional regularity conditions (A.3)(b) and (A.5)(b) and desirable properties of indicator and linear functions. The following theorem provides the consistency result of the variance estimator based on the efficient score function given in Theorem 2 . Denote P f = ∫ f ( y , δ , x , t ) d P ( y , δ , x , t ) and P n f = n − 1 ∑ i = 1 n f ( Y i , Δ i , X i , T i ) , where P is a probability measure, and P n is an empirical probability measure. Theorem 3 Suppose the conditions in Theorem 2 hold. Denote l ^ β ^ n ∗ ( Y , Δ , X , T ) = ∫ { X − X ‾ ( s ; β ^ n ) } { − ĝ ˙ n ( s ) } d M ^ ( s ) , where X ‾ ( s ; β ^ n ) = P n { X 1 ( Y − X T β ^ n ≥ s ) 1 ( T − X T β ^ n ≤ s ) } P n { 1 ( Y − X T β ^ n ≥ s ) 1 ( T − X T β ^ n ≤ s ) } and M ^ ( s ) = Δ 1 T − X T β ^ n ≤ Y − X T β ^ n ≤ s − ∫ − ∞ s 1 Y − X T β ^ n ≥ u 1 T − X T β ^ n ≤ u exp { ĝ n ( u ) } d u . Then P n l ^ β ^ n ∗ ( Y , Δ , X , T ) ⊗ 2 → P l β 0 ∗ ( Y , Δ , X , T ) ⊗ 2 in probability . In Theorem 3 , X ‾ ( s , β ^ n ) estimates E [ X | ϵ 0 ≥ s , τ 0 ≤ s ] from Theorem 2 . Once again, the proof of Theorem 3 follows Ding and Nan [ 25 ] with a major difference of including the truncation indicator, which can be easily dealt with. Due to a large duplication of Ding and Nan [ 19 ] in the proofs of above theorems, we omit all the details in this article. 4. A Simulation Study 4.1. Simulation Setup We perform an extensive simulation study to assess the validity of the proposed methodology. To compare with existing methods, we simulate datasets following the exact same setup as used in Ning et al. [ 12 ] from the following linear model: log ( Y ˜ ) = 1 + β 1 X 1 + β 2 X 2 + e , (8) where β 1 = 0 . 5 and β 2 = 1 . We generate binary covariate X 1 from a Bernoulli (0.5) distribution, continuous covariate X 2 from a Uniform (0,1) distribution, truncation time T ˜ from a Uniform (0, c ) distribution, censoring times C ˜ from a Uniform (0, c ) distribution, and the error term e from either a Uniform (−0.5, 0.5) distribution or a N ( 0 , 1 / 12 ) distribution. Then Y ˜ is determined from the above Equation ( 8 ). The constant c is chosen to give a 15%, or 30%, or 50% censoring rate. The simulated censoring and truncation distributions reflect a real study setup where censoring or study entry could occur at any point during the fixed duration of the study. Simulations are run with sample sizes of 100 and 200 at the three previously mentioned levels of censoring. Simulation results are directly compared to those reported in Ning et al. [ 12 ]. We further expand the simulation to demonstrate the robustness of the sieve likelihood method under different error distributions. Specifically, the error term comes from one of three distributions: (a) generalized extreme value distribution, or Gumbel distribution, (b) a 50% even mixture of N ( 0 , 1 / 24 ) and N ( 0 , 1 / 8 ) distributions, and (c) a 50% even mixture of N ( 0 , 1 / 12 ) and N ( − 0 . 3 , 1 / 24 ) distributions. These distributions were chosen to have variances the same as or similar to those in Ning et al. [ 12 ] while exhibiting more complex nuisance parameter curves. In both parts, the number of internal knots, K , used to estimate the nuisance parameter for each setup was determined via cross‐validation and ranged between 1 and 4. Since left‐truncation removes data, we keep simulating independent data until accumulated non‐truncated data reach the desired sample size, then treat the sample size as a known quantity. For each simulation setup, we generate 1000 datasets. We use 10 randomly selected initial values as described in Section 2 when fitting each model. 4.2. Simulation Results Table 1 shows a comparison of the results of the first part of our simulation study to those presented in Section 4 of Ning et al. [ 12 ]. The comparison methods are the semi‐rank‐based estimator proposed in Ning et al. [ 12 ], the inverse weighted estimating equation estimator proposed in Shen et al. [ 27 ], the Buckley‐James‐type estimator from Ning et al. [ 28 ], and the rank‐based estimator proposed by Lai and Yin [ 7 ]. Among these methods, , , and all use the length‐biased approach, meaning that they assume a uniform truncation distribution, whereas applies to general left‐truncated data. TABLE 1. Comparison of the Sieve Likelihood method to other approaches as reported in Ning et al. [ 12 ] in simulations. Results of other approaches are taken directly from Ning et al. [ 12 ], where is the estimator proposed in Ning et al. [ 12 ]; is the inverse weighted estimating equation from Shen et al. [ 27 ]; is the Buckley‐James‐type estimating equation from Ning et al. [ 28 ]; is the rank‐based approach presented in Lai and Ying [ 7 ]. For further details on these approaches see Ning et al. [ 12 ]. All variances and biases are multiplied by 1 0 3 . N Cens. % Sieve Likelihood β 1 β 2 β 1 β 2 β 1 β 2 β 1 β 2 β 1 β 2 Bias Bias Bias Var Bias Var Bias Var Bias Var Bias Var Bias Var Bias Var Bias Var e ∼ Unif ( − 0 . 5 , 0 . 5 ) 100 15% 1.6 1.4 5.2 5.0 −4.0 1.9 −6.0 6.2 −7.0 4.9 −5.0 14.9 0.0 4.6 2.0 2.6 1.0 8.5 −11.0 4.1 100 30% 2.2 2.1 0.7 7.5 8.0 2.6 27.0 9.0 −7.0 5.3 −6.0 18.5 −5.0 4.5 −1.0 14.4 0.0 3.5 −6.0 13.7 100 50% −0.7 3.1 9.5 11.1 −18.0 4.9 −33.0 16.1 −8.0 7.6 −5.0 24.7 15.0 11.3 36.0 36.0 0.0 6.1 −60.0 35.0 200 15% 1.3 0.6 −1.1 2.1 −3.0 0.7 −7.0 2.4 −3.0 2.5 0.0 7.6 1.0 2.3 −1.0 6.4 1.0 0.9 −1.0 2.9 200 30% −0.2 1.0 −1.9 3.5 −11.0 1.0 −25.0 3.9 −3.0 2.5 −5.0 8.8 0.0 2.3 2.0 7.4 1.0 1.3 −3.0 4.5 200 50% −3.6 1.6 0.5 5.5 −13.0 1.8 −28.0 6.1 0.0 3.4 0.0 11.4 15.0 4.4 36.0 14.4 1.0 2.1 −14.0 10.4 e ∼ N ( 0 , 1 / 12 ) 100 15% −0.5 4.1 8.8 13.4 −6.0 5.3 −19.0 14.6 −2.0 5.3 −11.0 15.4 −1.0 5.2 −2.0 13.9 4.0 7.4 −4.0 19.0 100 30% −1.4 4.4 6.5 14.6 −18.0 5.8 −43.0 18.5 −2.0 5.8 −13.0 18.0 1.0 4.5 −5.0 15.2 1.0 8.6 −14.0 22.8 100 50% 29.8 3331.6 4.0 18.6 −29.0 7.6 −44.0 26.2 −3.0 7.6 −12.0 25.6 −30.0 9.4 −65.0 34.6 −2.0 10.8 −60.0 33.5 200 15% −1.4 2.2 2.6 6.8 −10.0 2.4 −24.0 7.7 −2.0 2.6 −11.0 7.4 −1.0 2.4 −1.0 6.6 0.0 3.6 −3.0 9.4 200 30% 0.7 2.5 1.8 7.8 −21.0 2.6 −44.0 9.2 −2.0 2.9 −12.0 8.6 −1.0 2.5 2.0 7.2 −1.0 4.0 −6.0 11.2 200 50% −0.2 3.0 6.7 9.9 −28.0 3.6 −31.0 13.5 −2.0 3.5 −9.0 12.3 −30.0 4.2 −52.0 14.6 −3.0 4.9 −12.0 15.6 Open in a new tab The bias of our approach is similar to that of the other presented approaches, and in almost all cases the variance of our estimator is smaller than that of all comparison approaches presented in Ning et al. [ 12 ]. The one clear exception is in the 50% censoring case when the sample size is 100. This is expected, as the high censoring rate leads to a much smaller sample size, undermining the effectiveness of the method of sieves. In Ning et al. [ 12 ], the authors state that based on the outcome of their simulation study, there is “no uniform best estimation method in terms of statistical efficiency among the estimation methods considered.” We see that our sieve likelihood approach does indeed yield a uniformly most efficient estimator, as long as the number of observed events is large enough for splines to be used effectively. This aligns with the semiparametric efficiency theory from which the estimator was derived. Results from the second part of our simulation study are presented in Table 2 . For both β 1 and β 2 the biases are negligibly low, although the magnitude of the bias is often larger for β 2 which likely stems from the lower signal‐to‐noise ratio for that covariate. TABLE 2. Results from numerical simulations, where β 1 = 0 . 5 and β 2 = 1 . Error distributions: (a) Generalized Extreme Value Distribution, (b) 0 . 5 N ( 0 , 1 / 24 ) + 0 . 5 N ( 0 , 1 / 8 ) , (c) 0 . 5 N ( 0 , 1 / 12 ) + 0 . 5 N ( − 0 . 3 , 1 / 24 ) . Truncation probability ranged between 60% and 90%, depending on the error and censoring distributions. is based on the efficient score function estimator in Theorem 3 , is based on inverting the information matrix for β and γ , and is the empirical variance. The coverage probability is obtained using . Error Sample size Censoring % Bias × 1 0 3 95% Coverage β 1 β 2 β 1 β 2 β 1 β 2 β 1 β 2 β 1 β 2 (a) 200 15% 0.9 12.7 1.5 4.8 1.8 5.5 1.7 5.3 93.4 93.3 30% 0.5 13.4 1.6 5.1 1.9 6.1 1.7 6.1 93.8 91.5 50% 2.7 15.6 1.7 5.9 2.0 7.3 2.2 6.9 91.2 92.4 400 15% −0.1 10.1 0.8 2.5 0.8 2.6 1.0 2.9 93.1 93.2 30% 1.4 11.6 0.8 2.6 0.9 2.7 1.0 2.9 93.7 92.7 50% 5.7 13.9 0.9 2.9 1.0 3.1 1.0 3.3 93.4 92.8 800 15% −0.7 10.0 0.4 1.2 0.4 1.3 0.4 1.4 94.5 93.0 30% 1.3 9.4 0.4 1.3 0.4 1.3 0.4 1.5 94.6 93.2 50% 5.3 15.2 0.4 1.4 0.5 1.5 0.5 1.6 92.7 91.2 (b) 200 15% 1.5 9.1 2.1 6.6 2.5 8.6 2.8 7.9 91.5 92.3 30% 4.1 5.6 2.3 7.0 2.7 9.0 3.0 9.1 92.1 90.9 50% 0.4 2.7 2.6 8.4 3.4 11.6 3.6 10.8 92.6 92.2 400 15% −4.3 −5.8 1.0 3.2 1.2 4.0 1.3 4.2 93.1 92.2 30% −0.9 −4.9 1.1 3.3 1.2 4.0 1.4 4.2 92.4 92.4 50% −3.6 −12.3 1.2 3.9 1.4 4.9 1.7 5.4 92.0 91.4 800 15% −1.8 −6.0 0.5 1.6 0.5 1.7 0.6 1.9 93.2 93.9 30% −3.4 −8.7 0.6 1.7 0.6 2.0 0.7 2.2 93.1 92.2 50% −3.4 −9.6 0.6 2.0 0.7 2.3 0.8 2.7 92.1 91.0 (c) 200 15% 1.7 2.2 2.1 6.4 3.0 10.4 2.5 9.2 92.5 89.8 30% 2.0 3.3 2.2 6.9 4.4 13.3 3.0 9.5 91.2 90.3 50% 3.4 3.4 2.5 8.5 5.0 13.2 3.3 12.6 91.8 89.8 400 15% −0.2 −2.1 1.1 3.3 1.4 4.5 1.4 4.0 92.1 92.7 30% −1.3 −4.2 1.1 3.6 1.4 5.2 1.4 4.9 92.4 91.2 50% 0.0 −7.4 1.3 4.2 1.9 6.5 1.7 5.3 92.0 91.5 800 15% −0.1 −2.9 0.6 1.7 0.6 2.6 0.6 2.1 95.4 92.6 30% −0.2 −2.5 0.6 1.8 0.6 2.0 0.6 2.1 94.9 93.0 50% −1.0 −5.5 0.7 2.1 0.8 2.7 0.8 2.7 94.0 91.3 Open in a new tab The variance estimates across the three methods of estimation are very similar, especially at larger sample sizes. While the theory presented in this work only guarantees that the variance estimator obtained from the efficient score function converges to the truth, we see that the variance estimator obtained by inverting the information matrix of ( β , γ ) is close to the value of , and they both are close to the empirical variance . Further, all three variances are reduced approximately by half as the sample size is doubled, which is expected from the large sample theory. The confidence intervals created using provide proper coverages. The same pattern holds irrespective of error distributions. Recall that we use 10 random starting points to fit the model to each dataset, and select the result with the highest log‐likelihood value. Given the known truth, we can examine the convergence of these 10 random starts for each simulated dataset. For both β 1 and β 2 , we find about 80 ∼ 95 % of the randomly selected initial values lead to convergence somewhere close to the truth. Along with the regression coefficient estimates, we also examine the estimates of the infinite‐dimensional nuisance parameter, the logarithm of the hazard function, in simulations with a sample size of 800. Figure 1 summarizes the simulation results for each error distribution, including a sample of 100 individual estimates plotted with light‐blue solid lines, the point‐wise average over all 1000 datasets plotted with a red solid line, and the ground truth plotted with a black dashed line. The bounds of the x ‐axis represent the middle 95% of the theoretical error distribution. We see from Figure 1 that the spline estimates of log ( λ 0 ) perform reasonably well in simulations. FIGURE 1. Open in a new tab Plots showing the spline estimator for the log hazard function of each error distribution at 15%, 30%, and 50% censoring. A sample of individual estimators obtained from 100 datasets are plotted with light‐blue solid lines, along with the point‐wise average of the estimators obtained from all 1000 datasets plotted with red solid line and the true value plotted with black dashed line. Error distributions: (a) Generalized Extreme Value Distribution, (b) 0.5N(0,1/24) + 0.5N(0, 1/8), (c) 0.5N(0,1/12) + 0.5N(−0.3, 1/24). 5. Real Data Applications 5.1. The Canadian Study of Health and Aging The Canadian Study of Health and Aging (CSHA) is a multicenter epidemiologic study of dementia and related conditions among Canadians aged 65 years and older [ 20 , 21 ]. In the first phase (1991–1992), 10,263 participants were enrolled from a stratified random sample of community and institutional residents across Canada. Community‐dwelling participants completed a brief cognitive screening, and those with lower scores (together with a random sample of higher scorers) underwent a more detailed clinical assessment following common evaluation practices. Participants were classified into major dementia subtypes using standardized clinical guidelines. In the follow‐up phase (1996–1997), information on survival and date of death was obtained for all participants diagnosed with dementia at baseline. Because participants entered the study only after dementia onset had occurred, the data are left‐truncated, as individuals who developed dementia but died before enrollment were not observed. This means that the observed sample overrepresents longer survivors and underrepresents those with shorter post‐dementia lifespans. In addition, some subjects were still alive at the end of follow‐up, resulting in right‐censored observations. Fitting a standard survival model to the observed data would yield biased estimates of survival time and covariate effects due to the left‐truncation, underscoring the need for statistical methods that explicitly account for both truncation and censoring. The present analysis focuses on analyzing the time from dementia onset, the reference age A 0 , to death and parallels work done in Wolfson et al. [ 29 ] and Ning et al. [ 12 ] under the length‐biased sampling paradigm. After excluding subjects with missing onset dates or other dementia types, 821 individuals remained (395 probable Alzheimer's, 252 possible Alzheimer's, and 174 vascular dementia). Demographic information for each group is shown in Table 3 . TABLE 3. Descriptive statistics of the 821 patients analyzed from the CHSA data. Variable Mean (SD) or N (%) Probable Alzheimer's Possible Alzheimer's Vascular Dementia Age of Dementia Onset 82.3 (7.7) 81.0 (8.9) 78.5 (9.0) Survival (post‐diagnosis, in years) * 6.3 (3.9) 7.2 (5.6) 6.4 (6.4) Censored Patients 85 (21.5%) 60 (23.8%) 34 (19.5%) Sex Female 312 (79.0%) 171 (67.9%) 100 (57.5%) Male 83 (21.0%) 81 (32.1%) 74 (42.5%) Education (years) 8.6 (3.8) 8.6 (3.5) 8.6 (3.9) ≤ 8 years 166 (42.0%) 119 (47.2%) 83 (47.7%) > 8 years 149 (37.7%) 93 (36.9%) 67 (38.5%) Missing 80 (20.3%) 40 (15.9%) 24 (13.8%) Open in a new tab Abbreviation: SD, standard deviation. * Only computed from non‐censored subjects. We fit a semiparametric linear model using the sieve likelihood approach, adjusting for age of dementia onset and sex. To avoid dropping data (and since the values are similar across dementia types) we did not include education in our model. The form of the model is given in Equation ( 9 ) and parameter estimates are presented in Table 4 . Five‐fold cross‐validation was used to determine the optimal number of knots, which was one in the final model. log ( Time to Death (in days) ) = β 1 1 ( Diagnosis = Probable AD ) + β 2 1 ( Diagnosis = Vascular Dementia ) + β 3 ( Dementia Onset Age ) + β 4 1 ( Sex = Female ) + Error (9) TABLE 4. Results of the sieve method for fitting model ( 9 ) to the left‐truncated right‐censored CHSA data. Variable Estimate 95% CI P ‐value Type of Dementia Possible Alzheimer's Referent — — Probable Alzheimer's 0.094 ( − 0 . 043 , 0 . 231 ) 0.177 Vascular Dementia 0.055 ( − 0 . 063 , 0 . 175 ) 0.362 Age at Disease Onset 0.004 ( − 0 . 003 , 0 . 011 ) 0.342 Sex Male Referent — — Female 0.044 ( − 0 . 077 , 0 . 164 ) 0.477 Open in a new tab From the fitted results, we conclude that even after adjusting for age at disease onset and sex, there is no evidence of a difference in survival times between the different dementia categories considered here. As a comparison, we also ran an unadjusted analysis and obtained estimates within the confidence intervals estimated by Ning et al. [ 12 ]. In other words, we reached the same conclusions as in Ning et al. [ 12 ], where no difference was seen between the groups in the unadjusted analysis. 5.2. The 90+ Study of Dementia To further illustrate the proposed methodology, we also apply our method to the dataset of The 90+ Study of Alzheimer's disease and dementia [ 1 ]. This dataset comes from an ongoing study at the Alzheimer's Disease Research Center at the University of California, Irvine. The study enrolled participants who were at least 90 years old and did not have Alzheimer's disease or dementia at baseline. These individuals were followed via yearly visits. The main feature of interest is a diagnosis of the cognitive state of each participant, which comes from a panel of experts after reviewing the results of neurologic examination and a battery of neuropsychological tests. Understanding dementia onset during lifespan is of great interest. Because dementia may or may not occur during one's lifetime, a clear description of such a problem involves two components: the lifetime dementia‐free probability and the age distribution of dementia onset during one's lifetime given dementia has occurred. These two components can be modeled separately conditional on lifetime, where lifetime becomes a covariate in both conditional models. It is well‐known that complete‐case analysis when a covariate is subject to right censoring is a valid approach to estimate regression coefficients, see Kong et al. [ 30 , 31 ], and references therein. Thus we apply complete‐case analysis to the second component in the analysis of The 90+ Study where only individuals with available death time are included in the analysis. Specifically, we are interested in understanding the impact of demographic variables on the age of dementia onset relative to death time among participants who have developed dementia at some point during their lifetime by modeling the proportion of life after age 90 lived in a healthy state. For this purpose, we restrict our analysis to 349 available participants with complete information of age at death and age at dementia, out of the total of 921 patients available in the dataset. These original 921 participants were from two main cohorts recruited at two different time periods: The first cohort (599 out of 921 patients; 278 out of 349 in our sample of interest) is made up of survivors from the Leisure World Cohort Study (LCWS) and is primarily Caucasian, well‐educated, upper middle‐class, and mostly female [ 1 , 32 ]. The LCWS is a health survey study of the residents of the Leisure World (now known as Laguna Woods Village), a large retirement community in Orange County, California. This cohort was recruited into the study between 2003 and 2006. The second cohort (322 out of 921 patients; 71 out of 349 in our sample of interest) was recruited between 2014 and 2017, using a variety of methods including community outreach, earned media, direct mailings, and referrals [ 33 ]. The second cohort has a higher proportion of males, a higher proportion of subjects who graduate college, and similar rates of stroke and heart disease. Since participants enrolled in the study at different ages that range from 90.1 to 103.0 years old, the data we consider in this analysis are left‐truncated. We do not have right‐censoring in the dataset because ages at dementia onset are observed for all participants. Some details of the dataset are given in Table 5 , where age at dementia ranges from 90.9 to 109.8 years and age at death ranges from 91.3 to 111.2 years. TABLE 5. Descriptive statistics of the 349 participants of The 90+ Study. Variable Mean (SD) or N (%) Enrollment age 93.201 (2.745) Dementia diagnosis age 96.343 (3.222) Death age 98.514 (3.374) 0.737 (0.201) Cohort First 278 (79.7%) Second 71 (20.3%) Sex Female 251 (71.9%) Male 98 (28.1%) Education Did not graduate college 208 (59.6%) Graduated college 141 (40.4%) Has had a stroke? No 211 (60.5%) Yes 138 (39.5%) Has had heart disease? No 142 (40.7%) Yes 207 (59.3%) Open in a new tab Note: SD, standard deviation; , the proportion of lifespan after 90 years of age that is lived without dementia. To better answer the question of interest, we derive a new variable called the Healthspan Proportion for Age 90, HP 90 , that is the number of healthy years after the reference age A 0 = 90 divided by the number of total years lived after age 90. In other words, HP 90 is defined as [(Age at Dementia) − A 0 ]/[(Age at Death) − A 0 ]. In order to correctly fit the model, we must also scale the truncation time in this same way. Higher HP 90 implies better quality of life for an individual in this population of oldest old. We consider the following linear regression model for the logit transformed HP 90 and leave the error distribution unspecified: log HP 90 1 − HP 90 = β 1 ( Death Age − 90 ) + β 2 1 ( Cohort = 2 ) + β 3 1 ( Sex = Female ) + β 4 1 ( Education = Graduated College ) + β 5 1 ( Stroke = Yes or Heart Disease = Yes ) + Error . (10) Since HP 90 potentially ranges from 0 to 1, the logit transformed HP 90 can take any real value, and it is left‐truncated because age at dementia is left‐truncated. We apply our proposed method to fit a left‐truncated regression model to the data. We use five‐fold cross‐validation to select the number of internal knots for the considered B‐spline basis expansion, and ultimately choose two internal knots. From the regression results given in Table 6 we see that increased longevity is associated with higher HP 90 , and the second cohort has higher HP 90 than the first cohort. TABLE 6. Results of the sieve method for fitting model ( 10 ) to the left‐truncated data of The 90+ Study. Variable Estimate 95% CI P ‐value Death age 0.0578 (0.0221, 0.936) 0.0015 Cohort First Referent — — Second 0.3312 (0.0158, 0.6465) 0.0396 Sex Male Referent — — Female −0.0660 ( − 0 . 3431 , 0 . 2109 ) 0.6403 Education Did not graduate college Referent — — Graduated college 0.1638 ( − 0 . 0882 , 0 . 4157 ) 0.2026 Comorbidity: Stroke or CHD No Referent — — Yes 0.0200 ( − 0 . 3332 , 0 . 3733 ) 0.9116 Open in a new tab 6. Discussion Left‐truncation and right‐censoring are often present in prospective cohort studies. Although linear regression models with unspecified error distributions are appealing due to simplicity, their applications to censored or truncated event time data are rather limited, mainly because of lacking easily implementable estimating methods with desirable statistical properties. In this work we show that the semiparametric sieve maximum likelihood approach for bundled parameters possesses nice properties in dealing with left‐truncated and right‐censored data in linear regression models. The method yields asymptotically efficient and normally distributed estimates for regression coefficients, with estimators obtained using an easily implementable gradient search algorithm. We expect this work would encourage the proper use of linear regression models in analyzing event time data that are subject to left‐truncation and right‐censoring. Funding This work was supported by the National Institutes of Health (Grant Nos. RF1 AG075107 and P30 AG066519) and the National Science Foundation (Grant No. DMS 2412746). Conflicts of Interest The authors declare no conflicts of interest. Acknowledgments We are grateful to Dr. Maria Corrada for sharing The 90+ Study dataset. The CSHA was supported by the Seniors Independence Research Program, through the National Health Research and Development Program (NHRDP) of Health Canada (project 6606‐3954‐MC[S]). The progression of dementia project within the CSHA was supported by Pfizer Canada through the Health Activity Program of the Medical Research Council of Canada and the Pharmaceutical Manufacturers Association of Canada; by the NHRDP (project 6603‐1417‐302[RJ]); by Bayer; and by the British Columbia Health Research Foundation (projects 38 [93‐2] and 34 [96‐1]). Appendix A. Likelihood Function Let F ‾ A ( a ) denote the survival function of an arbitrary random variable A . We proceed with looking at two sub‐distributions that correspond to Δ = 1 and Δ = 0 , respectively. We have pr ( Y ≤ y , Δ = δ | Y > T , T = t , X = x ) = pr ( Y ≤ y , Δ = δ , Y > t | T = t , X = x ) pr ( Y > t | T = t , X = x ) 1 ( y > t ) . (A1) From our assumption, Y ∗ is independent of ( C , T ) given X . Consider first the denominator in ( A1 ). We have pr ( Y > t | T = t , X = x ) = pr ( Y ∗ > t , C > t | T = t , X = x ) = F ‾ Y ∗ | X ( t | x ) F ‾ C | T , X ( t | t , x ) . (A2) Now consider the numerator in ( A1 ). For the case where Δ = 1 , we have pr ( Y ≤ y , Δ = 1 , Y > t | T = t , X = x ) = pr ( Y ∗ ≤ y , Y ∗ < C , Y ∗ > t | T = t , X = x ) = ∫ t y ∫ u ∞ f Y ∗ , C | T , X ( u , v | t , x ) d v d u = ∫ t y ∫ u ∞ f Y ∗ | X ( u | x ) f C | T , X ( v | t , x ) d v d u = ∫ t y f Y ∗ | X ( u | x ) F ‾ C | T , X ( u | t , x ) d u . When Δ = 0 we have pr ( Y ≤ y , Δ = 0 , Y > t | T = t , X = x ) = pr ( C ≤ y , Y ∗ > C , C > t | T = t , X = x ) = ∫ t y ∫ v ∞ f Y ∗ , C | T , X ( u , v | t , x ) d u d v = ∫ t y ∫ v ∞ f Y ∗ | X ( u | x ) f C | T , X ( v | t , x ) d u d v = ∫ t y f C | T , X ( v | t , x ) F ‾ Y ∗ | X ( v | x ) d v . Thus we obtain the following conditional density function: f Y , Δ | Y > T , T , X ( y , δ | y > t , t , x ) = f Y ∗ | X ( y | x ) F ‾ C | T , X ( y | t , x ) δ f C | T , X ( y | t , x ) F ‾ Y ∗ | X ( y | x ) 1 − δ F ‾ Y ∗ | X ( t | x ) F ‾ C | T , X ( t | t , x ) 1 ( y > t ) . Then the joint density function of left‐truncated data is f Y , Δ , T , X | Y > T ( y , δ , t , x | y > t ) 1 ( y > t ) = f Y , Δ | T , X , Y > T ( y , δ | t , x , y > t ) f T , X | Y > T ( t , x | y > t ) 1 ( y > t ) = f Y ∗ | X ( y | x ) F ‾ C | T , X ( y | t , x ) δ f C | T , X ( y | t , x ) F ‾ Y ∗ | X ( y | x ) 1 − δ F ‾ Y ∗ | X ( t | x ) F ‾ C | T , X ( t | t , x ) × f T , X | Y > T ( t , x | y > t ) 1 ( y > t ) . (A3) Dropping the factors that are free of ( β , λ e ) from ( A3 ), we obtain the likelihood: L ( β , λ e | Y , Δ , T , X ) = f Y ∗ | X ( Y | X ) Δ F ‾ Y ∗ | X ( Y | X ) 1 − Δ F ‾ Y ∗ | X ( T | X ) . Using hazard functions, we can rewrite the above likelihood function as the following: L ( β , λ e | Y , Δ , T , X ) = λ Y ∗ | X ( Y | X ) Δ exp − ∫ − ∞ Y λ Y ∗ | X ( u | X ) d u exp − ∫ − ∞ T λ Y ∗ | X ( u | X ) d u = λ Y ∗ | X ( Y | X ) Δ exp − ∫ T Y λ Y ∗ | X ( u | X ) d u = λ e ( Y − X T β ) Δ exp − ∫ T − X T β Y − X T β λ e ( u ) d u . Data Availability Statement The data that support the findings of this study are available on request from the corresponding author. The data are not publicly available due to privacy or ethical restrictions. References 1. Corrada M. M., Brookmeyer R., Paganini‐Hill A., Berlau D., and Kawas C. H., “Dementia Incidence Continues to Increase With Age in the Oldest Old: The 90+ Study,” Annals of Neurology 67, no. 1 (2010): 114–121. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 2. Chen L. P. and Yi G. Y., “Semiparametric Methods for Left‐Truncated and Right‐Censored Survival Data With Covariate Measurement Error,” Annals of the Institute of Statistical Mathematics 73, no. 3 (2021): 481–517. [ Google Scholar ] 3. Kalbfleisch J. D. and Prentice R. L., The Statistical Analysis of Failure Time Data, 2nd ed. (J. Wiley, 2002). [ Google Scholar ] 4. Wang X. and Wang Q., “Semiparametric Linear Transformation Model With Differential Measurement Error and Validation Sampling,” Journal of Multivariate Analysis 141 (2015): 67–80. [ Google Scholar ] 5. Adichie J. N., “Estimates of Regression Parameters Based on Rank Tests,” Annals of Mathematical Statistics 38, no. 3 (1967): 894–904. [ Google Scholar ] 6. Bhattacharya P. K., Chernoff H., and Yang S. S., “Nonparametric Estimation of the Slope of a Truncated Regression,” Annals of Statistics 11, no. 2 (1983): 505–514. [ Google Scholar ] 7. Lai T. L. and Ying Z., “Rank Regression Methods for Left‐Truncated and Right‐Censored Data,” Annals of Statistics 19, no. 2 (1991): 531–556. [ Google Scholar ] 8. Lai T. L. and Ying Z., “Linear Rank Statistics in Regression Analysis With Censored or Truncated Data,” Journal of Multivariate Analysis 40, no. 1 (1992): 13–45. [ Google Scholar ] 9. Lai T. L. and Ying Z., “A Missing Information Principle and M‐Estimators in Regression Analysis With Censored and Truncated Data,” Annals of Statistics 22, no. 3 (1994): 1222–1255. [ Google Scholar ] 10. Kim C. K. and Lai T. L., “Efficient Score Estimation and Adaptive M‐Estimators in Censored and Truncated Regression Models,” Statistica Sinica 10, no. 3 (2000): 731–749. [ Google Scholar ] 11. Chen L. P. and Qiu B., “Analysis of Length‐Biased and Partly Interval‐Censored Survival Data With Mismeasured Covariates,” Biometrics 79, no. 4 (2023): 3929–3940. [ DOI ] [ PubMed ] [ Google Scholar ] 12. Ning J., Qin J., and Yu Y., “Semiparametric Accelerated Failure Time Model for Length‐Biased Data With Application to Dementia Study,” Statistica Sinica 24, no. 1 (2014): 313–333. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 13. Chen L. P., “Feature Screening via Concordance Indices for Left‐Truncated and Right‐Censored Survival Data,” Journal of Statistical Planning and Inference 232 (2024): 106153. [ Google Scholar ] 14. Wang M. C., “A Semiparametric Model for Randomly Truncated Data,” Journal of the American Statistical Association 84, no. 407 (1989): 742–748. [ Google Scholar ] 15. Geman S. and Hwang C. R., “Nonparametric Maximum Likelihood Estimation by the Method of Sieves,” Annals of Statistics 10, no. 2 (1982): 401–414. [ Google Scholar ] 16. Shen X. and Wong W. H., “Convergence Rate of Sieve Estimates,” Annals of Statistics 22, no. 2 (1994): 580–615. [ Google Scholar ] 17. [17] Chen X., “Large Sample Sieve Estimation of Semi‐Nonparametric Models,” in Handbook of Econometrics, vol. 6, ed. Heckman J. J. and Leamer E. E. (Elsevier, 2007), 5549–5632. [ Google Scholar ] 18. van der Vaart A. W. and Wellner J. A., Weak Convergence and Empirical Processes: With Applications to Statistics, 2nd ed. (Springer International Publishing AG, 2023). [ Google Scholar ] 19. Ding Y. and Nan B., “A Sieve M‐Theorem for Bundled Parameters in Semiparametric Models, With Application to the Efficient Estimation in a Linear Model for Censored Data,” Annals of Statistics 39, no. 6 (2011): 2795–3443. [ PMC free article ] [ PubMed ] [ Google Scholar ] 20. The Canadian Study of Health and Aging Working Group , “Canadian Study of Health and Aging: Study Methods and Prevalence of Dementia,” Canadian Medical Association Journal 150, no. 6 (1994): 899–913. [ PMC free article ] [ PubMed ] [ Google Scholar ] 21. The Canadian Study of Health and Aging Working Group , “The Canadian Study of Health and Aging: Risk Factors for Alzheimer's Disease in Canada,” Neurology 44, no. 11 (1994): 2073. [ DOI ] [ PubMed ] [ Google Scholar ] 22. De Boor C., A Practical Guide to Splines, rev. ed. (Springer, 2001). [ Google Scholar ] 23. Lai T. L. and Ying Z., “Asymptotically Efficient Estimation in Censored and Truncated Regression Models,” Statistica Sinica 2, no. 1 (1992): 17–46. [ Google Scholar ] 24. Schumaker L. L., Spline Functions: Basic Theory, 3rd ed. (Cambridge University Press, 2007). [ Google Scholar ] 25. Ding Y. and Nan B., “Supplement to “A Sieve M‐Theorem for Bundled Parameters in Semiparametric Models, With Application to the Efficient Estimation in a Linear Model for Censored Data”,” Annals of Statistics 39, no. 6 (2011): Supp 1–13. [ PMC free article ] [ PubMed ] [ Google Scholar ] 26. Zhao X., Wu Y., and Yin G., “Sieve Maximum Likelihood Estimation for a General Class of Accelerated Hazards Models With Bundled Parameters,” Bernoulli 23, no. 4B (2017): 3385–3411. [ Google Scholar ] 27. Shen Y., Ning J., and Qin J., “Analyzing Length‐Biased Data With Semiparametric Transformation and Accelerated Failure Time Models,” Journal of the American Statistical Association 104, no. 487 (2009): 1192–1202. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 28. Ning J., Qin J., and Shen Y., “Buckley‐James‐Type Estimator With Right‐Censored and Length‐Biased Data,” Biometrics 67, no. 4 (2011): 1369–1378. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 29. Wolfson C., Wolfson D. B., Asgharian M., et al., “A Reevaluation of the Duration of Survival After the Onset of Dementia,” New England Journal of Medicine 344, no. 15 (2001): 1111–1116. [ DOI ] [ PubMed ] [ Google Scholar ] 30. Kong S. and Nan B., “Semiparametric Approach to Regression With a Covariate Subject to a Detection Limit,” Biometrika 103, no. 1 (2016): 161–174. [ Google Scholar ] 31. Kong S., Nan B., Kalbfleisch J., Saran R., and Hirth R., “Conditional Modeling of Longitudinal Data With Terminal Event,” Journal of the American Statistical Association 113, no. 521 (2018): 357–368. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 32. Corrada M. M., Brookmeyer R., Berlau D., Paganini‐Hill A., and Kawas C. H., “Prevalence of Dementia After Age 90: Results From the 90+ Study,” Neurology 71, no. 5 (2008): 337–343. [ DOI ] [ PubMed ] [ Google Scholar ] 33. Melikyan Z. A., Greenia D. E., Corrada M. M., Hester M. M., Kawas C. H., and Grill J. D., “Recruiting the Oldest‐Old for Clinical Research,” Alzheimer Disease and Associated Disorders 33, no. 2 (2019): 160–162. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Associated Data This section collects any data citations, data availability statements, or supplementary materials included in this article. Data Availability Statement The data that support the findings of this study are available on request from the corresponding author. The data are not publicly available due to privacy or ethical restrictions. Articles from Statistics in Medicine are provided here courtesy of Wiley ACTIONS View on publisher site PDF (547.5 KB) Cite Collections Permalink PERMALINK Copy RESOURCES Similar articles Cited by other articles Links to NCBI Databases Cite Copy Download .nbib .nbib Format: AMA APA MLA NLM Add to Collections Create a new collection Add to an existing collection Name your collection * Choose a collection Unable to load your collection due to an error Please try again Add Cancel Follow NCBI NCBI on X (formerly known as Twitter) NCBI on Facebook NCBI on LinkedIn NCBI on GitHub NCBI RSS feed Connect with NLM NLM on X (formerly known as Twitter) NLM on Facebook NLM on YouTube National Library of Medicine 8600 Rockville Pike Bethesda, MD 20894 Web Policies FOIA HHS Vulnerability Disclosure Help Accessibility Careers NLM NIH HHS USA.gov Back to Top

Record · ID 67529 · SHA-256 4424cd707ca995f4
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.