0% found this document useful (0 votes)

4 views

L6 Decision Tree Classifier

Uploaded by

Fahim Ahmed

We take content rights seriously. If you suspect this is your content, claim it here.

Available Formats

Download as PDF, TXT or read online on Scribd

0% found this document useful (0 votes)

4 views

L6 Decision Tree Classifier

Uploaded by

Fahim Ahmed

We take content rights seriously. If you suspect this is your content, claim it here.

Available Formats

Download as PDF, TXT or read online on Scribd

You are on page 1/ 46

Decision Trees

Course Title: Machine Learning

Dept. of Computer Science

Faculty of Science and Technology

Lecturer No: Week No: Semester: Summer 2021-22

Lecturer: Dr. M M Manjurul Islam
Overview

 What is a Decision Tree

 Sample Decision Trees

 How to Construct a Decision Tree

 Problems with Decision Trees

 Summary
Classification: Definition

 Given a collection of records (training set )

 Each record contains a set of attributes, one of the attributes is the
class.

 Find a model for class attribute as a function of the values

of other attributes.
 Goal: previously unseen records should be assigned a
class as accurately as possible.
 A test set is used to determine the accuracy of the model.
Usually, the given data set is divided into training and test sets,
with training set used to build the model and test set used to
validate it.
Illustrating Classification Task

Tid Attrib1 Attrib2 Attrib3 Class Learning

1 Yes Large 125K No
algorithm
2 No Medium 100K No

3 No Small 70K No

4 Yes Medium 120K No

Induction
5 No Large 95K Yes

6 No Medium 60K No

7 Yes Large 220K No Learn

8 No Small 85K Yes Model
9 No Medium 75K No

10 No Small 90K Yes

Model
10

Training Set
Apply
Tid Attrib1 Attrib2 Attrib3 Class Model
11 No Small 55K ?

12 Yes Medium 80K ?

13 Yes Large 110K ? Deduction

14 No Small 95K ?

15 No Large 67K ?
10

Test Set
What is a Decision Tree?

 An inductive learning task

 Use particular facts to make more generalized
conclusions

 A predictive model based on a branching series

of Boolean tests
 These smaller Boolean tests are less complex than a
one-stage classifier

 Let’s look at a sample decision tree…

Predicting Commute Time

If we leave at
Leave At 10 AM and
10 AM 9 AM
there are no
8 AM cars stalled on
Stall? Accident? the road, what
No Yes Long will our
No Yes
commute time
Short Long Medium Long be?
Inductive Learning

 In this decision tree, we made a series of

Boolean decisions and followed the
corresponding branch
 Did we leave at 10 AM?
 Did a car stall on the road?
 Is there an accident on the road?

 By answering each of these yes/no questions,

we then came to a conclusion on how long our
commute might take
Decision Trees as Rules

 We did not have represent this tree graphically

 We could have represented as a set of rules.

However, this may be much harder to read…
Decision Tree as a Rule Set
How to Create a Decision Tree

 We first make a list of attributes that we can

measure
 These attributes (for now) must be discrete

 We then choose a target attribute that we want

to predict
 Then create an experience table that lists what we
have seen in the past
Sample Experience Table

Example Attributes Target

Hour Weather Accident Stall Commute
D1 8 AM Sunny No No Long
D2 8 AM Cloudy No Yes Long
D3 10 AM Sunny No No Short
D4 9 AM Rainy Yes No Long
D5 9 AM Sunny Yes Yes Long
D6 10 AM Sunny No No Short
D7 10 AM Cloudy No No Short
D8 9 AM Rainy No No Medium
D9 9 AM Sunny Yes No Long
D10 10 AM Cloudy Yes Yes Long
D11 10 AM Rainy No No Short
D12 8 AM Cloudy Yes No Long
D13 9 AM Sunny No No Medium
Example of a Decision Tree

Splitting Attributes
Tid Refund Marital Taxable
Status Income Cheat

1 Yes Single 125K No Refund

2 No Married 100K No Yes No
3 No Single 70K No
NO MarSt
4 Yes Married 120K No
Single, Divorced Married
5 No Divorced 95K Yes
6 No Married 60K No TaxInc NO
7 Yes Divorced 220K No < 80K > 80K
8 No Single 85K Yes
NO YES
9 No Married 75K No
10 No Single 90K Yes
10

Model: Decision Tree

Training Data
Another Example of Decision Tree
MarSt Single,
Married Divorced
Tid Refund Marital Taxable
Status Income Cheat NO Refund
Yes No
1 Yes Single 125K No
2 No Married 100K No
NO TaxInc
3 No Single 70K No
< 80K > 80K
4 Yes Married 120K No
5 No Divorced 95K Yes NO YES
6 No Married 60K No
7 Yes Divorced 220K No
8 No Single 85K Yes There could be more than one tree that
9 No Married 75K No fits the same data!
10 No Single 90K Yes
10
Decision Tree Classification Task

Tid Attrib1 Attrib2 Attrib3 Class

Tree
1 Yes Large 125K No Induction
2 No Medium 100K No algorithm
3 No Small 70K No

4 Yes Medium 120K No

Induction
5 No Large 95K Yes

6 No Medium 60K No

7 Yes Large 220K No Learn

8 No Small 85K Yes Model
9 No Medium 75K No

10 No Small 90K Yes

Model
10

Training Set
Apply Decision
Model Tree
Tid Attrib1 Attrib2 Attrib3 Class
11 No Small 55K ?

12 Yes Medium 80K ?

13 Yes Large 110K ?

Deduction
14 No Small 95K ?

15 No Large 67K ?
10

Test Set
Apply Model to Test Data
Test Data
Start from the root of tree. Refund Marital Taxable
Status Income Cheat

No Married 80K ?
Refund 10

Yes No

NO MarSt
Single, Divorced Married

TaxInc NO
< 80K > 80K