Abstract: Processing a sparsely populated data source comprising: retrieving data from a sparsely populated data source (a plurality of records) to form a base dataset, each record comprising at least one unpopulated data field corresponding to a medical measurement; dividing the retrieved data into two portions, a training dataset and validation dataset; analyzing the training data set to obtain a trained model and measurement prediction protocols; imputing the predicted measurement values in the records of the training dataset; analyzing the training dataset; validating the disease model, wherein the records of the validation dataset comprise disease data associated with patient data, and determining a validation error; repeating steps to minimize the validation error and computing a prediction of a probable disease state for each patient record in the base dataset.