Question: KINDLY ANSWER USING PYTHON ONLY 9 . 2 Predicting Delayed Flights. The file FlightDelays.csv contains information on all commercial flights departing the Washington, DC area

KINDLY ANSWER USING PYTHON ONLY
9.2 Predicting Delayed Flights. The file FlightDelays.csv contains information on all commercial flights departing the Washington, DC area and arriving at New York during January 2004. For each flight, there is information on the departure and arrival airports, the distance of the route, the scheduled time and date of the flight, and so on. The variable that we are trying to predict is whether or not a flight is delayed. A delay is defined as an arrival that is at least 15 minutes later than scheduled.
Data Preprocessing. Transform variable day of week (DAY_WEEK) info a cate- gorical variable. Bin the scheduled departure time into eight bins. Use these and all
other columns as predictors (excluding DAY_OF_MONTH). Partition the data into training (60%) and validation (40%) sets.
a. Fit a classification tree to the flight delay variable using all the relevant predictors. Do not include DEP_TIME (actual departure time) in the model because it is unknown at the time of prediction (unless we are generating our predictions of delays after the plane takes off, which is unlikely). Use a tree with maximum depth 8 and minimum impurity decrease =0.01. Express the resulting tree as a set of rules.
b. If you needed to fly between DCA and EWR on a Monday at 7:00 AM, would you be able to use this tree? What other information would you need? Is it available in practice? What information is redundant?
c. Fit the same tree as in (a), this time excluding the Weather predictor. Display both the resulting (small) tree and the full-grown tree. You will find that the small tree contains a single terminal node.
i. Howisthesmalltreeusedforclassification?(Whatistheruleforclassifying?)
ii. Towhatisthisruleequivalent?
iii. Examine the full-grown tree. What are the top three predictors according to this tree?
iv. Why,technically,doesthesmalltreeresultinasinglenode?
v. Whatisthedisadvantageofusingthetoplevelsofthefull-growntreeasopposed
to the small tree?
vi. Compare this general result to that from logistic regression in the example in Chapter 10. What are possible reasons for the classification trees failure to find a good predictive model?

Step by Step Solution

There are 3 Steps involved in it

1 Expert Approved Answer
Step: 1 Unlock blur-text-image
Question Has Been Solved by an Expert!

Get step-by-step solutions from verified subject matter experts

Step: 2 Unlock
Step: 3 Unlock

Students Have Also Explored These Related Databases Questions!