Startup Survival & Series A Prediction

Predict whether a startup will raise another round within 36 months, and estimate Series A probability and pre-valuation.

Kiota IntelligenceVC AnalyticsR · SurvivalXGBoost

TL;DR: Clean labels and validation that match deployment prevalence make probability scores actionable. Prioritize label hygiene and calibration before optimizing model complexity.

The Labeling Challenge

In a Crunchbase dataset (~50k funding rounds), the naive label 'raised again = yes / no' is often misleading: many companies simply haven't had enough observable time to show a subsequent round. This is a label-validity problem, not a modeling one.

Two design decisions were essential:

  • Eligibility window: only include companies with at least the prediction horizon (36+ months) of observable history to avoid mislabeling recent companies as negatives.
  • Asymmetric holdout: set validation prevalence to match deployment conditions (not an artificial 50/50 split) so that probability calibration reflects the real world.

Models

Survival ensemble

A survival modeling stack (Cox baseline + gradient-boosted survival ensemble) estimates hazard over time and produces time-dependent probabilities of future fundraising. Useful to rank risk and expected time-to-event.

Calibrated classifier

An XGBoost/LightGBM classifier trained on eligible companies, then calibrated with isotonic regression/platt scaling to produce well-calibrated Series A probabilities for decision-making and scoring.

Pre-valuation ensemble

Stacked regression ensemble that combines gradient boosting, regularized linear models, and tree-based learners to predict pre-valuation figures (log-transformed). Outputs were used for valuation buckets in dashboards.

Impact

By improving label hygiene and aligning validation with deployment prevalence, the models provided stable probability scores used in investor filtering dashboards and early-warning monitoring. The focus on calibration and operational thresholds made the predictions actionable for investment teams.

Key Takeaway

Design insight: Clean labels and validation that match deployment prevalence make probability scores actionable. Prioritize label hygiene and calibration before optimizing model complexity.

FAQ

What is the key takeaway from "Startup Survival & Series A Prediction"?

Clean labels and validation that match deployment prevalence make probability scores actionable. Prioritize label hygiene and calibration before optimizing model complexity.

Who wrote this and what is it about?

This was written by Mahmoud Trigui, Senior Data Scientist. VC analytics: startup survival modeling, Series A probability, and pre-valuation estimation. Labeling and evaluation discussion.

Comments

Machine LearningFeature EngineeringMLForecastTime Series DecompositionForecastingLightGBMXGBoostCatboostClusteringSegmentationNLPLLMsWeb AppR MarkdownSQLOracle DBSAS-GuideSAS E-MinerDataikuBigQueryGCPPythonRCRISP-DMHypothesis TestingANOVAData AnalyticsDimensionality ReductionRecommendation SystemNetwork AnalysisGeospace AnalysisSpatial DataEmbeddingSampling TechniquesDecision RulesData StorytellingCVMChurnFraud DetectionSentiment AnalysisTopic ModelingIBM WatsonPowerBILooker StudioVBAStatistical LearningEnsemble ModelingStackingCross-ValidationProfilingABT ConstructionPlumberTidyverseShinyProphetDeep LearningScikit-LearnJSONSAS ProgrammingGitVS CodeCSS StylingAutomated ReportingOutlier DetectionTemporal ClusteringStartup SurvivalPre-Valuation ModelingK-MeansDecision TreesData SciencePredictive ModelingSVMLDAText ClassificationWeight PredictionPattern RecognitionReal-Time DetectionCommunity DetectionPipeline AutomationData Quality ChecksData ReliabilitySpecification MappingBusiness StrategyMarketing CampaignsTry & Buy FrameworksKPI DashboardsNetwork QualitySales AnalyticsMentoringStatistics LecturerRemote WorkHybrid WorkConsultingContractFull-TimeFreelanceSofrecomOrange GroupTunisia TelecomKiota IntelligenceVC AnalyticsSeries A PredictionProduction MLApplied AIPrompt EngineeringBusiness ForecastingDecision SystemsGraph AnalyticsHousehold DetectionMulti-SIM DetectionFTTH ForecastingAudit ExtractionInfrastructure ClassificationPydanticGPT-4OpenAI APIBase64 ClassificationZindiCodementorLAAS-CNRSESSAIMIT xPROTunisiaML CompetitionCell Tower AnalysisUber LogisticsUber Cape TownNecessary Condition AnalysisBehavioral SignalsSpike SmoothingObservation Unit DesignDendrogramWard ClusteringVIFTarget Encoding Machine LearningFeature EngineeringMLForecastTime Series DecompositionForecastingLightGBMXGBoostCatboostClusteringSegmentationNLPLLMsWeb AppR MarkdownSQLOracle DBSAS-GuideSAS E-MinerDataikuBigQueryGCPPythonRCRISP-DMHypothesis TestingANOVAData AnalyticsDimensionality ReductionRecommendation SystemNetwork AnalysisGeospace AnalysisSpatial DataEmbeddingSampling TechniquesDecision RulesData StorytellingCVMChurnFraud DetectionSentiment AnalysisTopic ModelingIBM WatsonPowerBILooker StudioVBAStatistical LearningEnsemble ModelingStackingCross-ValidationProfilingABT ConstructionPlumberTidyverseShinyProphetDeep LearningScikit-LearnJSONSAS ProgrammingGitVS CodeCSS StylingAutomated ReportingOutlier DetectionTemporal ClusteringStartup SurvivalPre-Valuation ModelingK-MeansDecision TreesData SciencePredictive ModelingSVMLDAText ClassificationWeight PredictionPattern RecognitionReal-Time DetectionCommunity DetectionPipeline AutomationData Quality ChecksData ReliabilitySpecification MappingBusiness StrategyMarketing CampaignsTry & Buy FrameworksKPI DashboardsNetwork QualitySales AnalyticsMentoringStatistics LecturerRemote WorkHybrid WorkConsultingContractFull-TimeFreelanceSofrecomOrange GroupTunisia TelecomKiota IntelligenceVC AnalyticsSeries A PredictionProduction MLApplied AIPrompt EngineeringBusiness ForecastingDecision SystemsGraph AnalyticsHousehold DetectionMulti-SIM DetectionFTTH ForecastingAudit ExtractionInfrastructure ClassificationPydanticGPT-4OpenAI APIBase64 ClassificationZindiCodementorLAAS-CNRSESSAIMIT xPROTunisiaML CompetitionCell Tower AnalysisUber LogisticsUber Cape TownNecessary Condition AnalysisBehavioral SignalsSpike SmoothingObservation Unit DesignDendrogramWard ClusteringVIFTarget Encoding