Machine learning and deep learning approaches are increasingly used to forecast long-term disease trajectories in individuals with chronic obstructive pulmonary disease (COPD). This study sought to synthesize the effectiveness of these prognostic tools for COPD, evaluate their performance relative to one another, and pinpoint major areas needing further investigation.We conducted a systematic review and meta-analysis to assess the effectiveness of machine learning- and deep learning-based prognostic models and to outline directions for future studies. We systematically searched PubMed, Embase, the Cochrane Library, ProQuest, Scopus, and Web of Science from their launch until April 6, 2023, targeting English-language publications that used machine learning or deep learning techniques to predict patient outcomes at least 6 months after the first clinical encounter in people diagnosed with COPD. Eligible studies included adults aged 18 to 90 years and accepted any input data. We presented the area under the receiver operating characteristic curve (AUC) along with 95% confidence intervals (CI) for forecasts of death, exacerbations, and reductions in forced expiratory volume in 1 s (FEV1). Heterogeneity across studies was quantified via Cochran’s Q test, where notable heterogeneity was indicated by P ≤ 0.10 or I² > 50%. We evaluated reporting standards with the TRIPOD checklist and conducted risk-of-bias evaluations using the PROBAST tool. The review was prospectively registered on PROSPERO (CRD42022323052). The initial literature search yielded 3620 records. After screening, 18 studies met the inclusion standards, including 12 that relied on traditional machine learning and 6 that utilized deep learning. Seven models focused on exacerbation probability, though only six provided AUC values with 95% CI from internal validation sets (combined AUC 0.77 [95% CI: 0.69–0.85]), accompanied by substantial heterogeneity (I² 97%, P < 0.0001). Eleven models addressed mortality risk, with six reporting AUC and 95% CI on internal validation data (pooled AUC 0.77 [95% CI: 0.74–0.80]) and moderate heterogeneity (I² 60%, P = 0.027). Two investigations examined declines in lung function but could not be aggregated. Machine learning and deep learning models did not show meaningful gains compared with established disease severity indices for exacerbation prediction (P = 0.24). Three studies offered head-to-head comparisons of machine learning against conventional severity scores for mortality, revealing no difference in pooled results (P = 0.57). Among the five studies conducting external validation, model performance was similar to or inferior to standard regression approaches. Key sources of bias included improper management of missing values, omission of model uncertainty estimates, and reliance on datasets undersized relative to the number of predictor variables. Available data provide scant support that conventional machine learning or deep learning prognostic models outperform existing disease severity scoring systems. Stronger compliance with established reporting standards in future work would reduce bias risks and improve replicability.