1 of 50

Project

BUFN 742�Financial Engineering

2022 FALL

TEAM 6

2 of 50

Project Deliverable

3 of 50

Fannie CAS Pricing: Fannie Mae CAS 2021 – RO1

​

We have used survival rate (cloglog function) model to predict the loss rate distribution of 10,000 loans via 500 paths of stochastic mortgage rates and House Prices.

​

We have successfully priced 6 CAS tranches, achieving the ideal results as follows:

​

​

4 of 50

Loss Distribution

6 Tranches from Fannie Mae;

Loss subordination:

(Class 1A-H:2%; Class 1M-1: 1.6%;

Class 1M-2: 1.25% ;Class IB-1: 0.7%;

Class 1B-2: 0.25% ;Class 1B-3H:0% )

Loss rate distribution:

​

  • The loss rate with the highest frequency is about 1%.
  • The far tail part has a loss rate about 2.25%;

5 of 50

Loss Distribution

Loss Rate Distribution & Tranches Loss Subordination

6 of 50

CRT Tranche Price and Yields

  • The senior tranche has the lowest yield about the risk free rate-3.39%

​

  • The junior tranche has the highest yield since its price is the lowest and it undertakes the highest loss risk;

​

  • Payoff frequency is used
  • The Ideal Pricing and Yields Result:

7 of 50

Sample Selection�

Week

1

8 of 50

Loan Sample SAS Code

9 of 50

Data Analysis�

Week

2

10 of 50

Distributional characteristics

 

  • After carefully observing the data, we noticed all the variables we chose have fixed values for the duration of the loan life. To avoid double counting when constructing the distribution charts, we only obtained the last recorded entry of each unique loan ID and used those numbers to generate the univariate and bivariate analysis. The table above shows the distributional characteristics of the chosen variables.

11 of 50

Univariate analysis 

State

Frequency of Categorical Variables

  • The univariate analysis allow insights into the distributional properties of each variable with respect to the total number of loans
  • We observe that a majority of the loans are originated from the state of California, taking up approximately 14.2% of all the loans

12 of 50

Univariate analysis 

Loan Purpose

Property Type

Occupancy Status

Frequency of Categorical Variables

13 of 50

Bivariate Analysis

  • To conduct bivariate analysis, we have to bin the continuous variables into different groups to show distributional properties of each group with regard to the number of defaults.
  • For simplicity, we group the continuous variables by their median values, we will split the data into more groups if they prove beneficial to our model.
  • The chart on the right shows the two FICO score groups, above or below 740, and their portion of defaults. The table above also shows some important statistics. 12581 loans defaulted in the group where FICO is above 740, and it makes up 3.73% of all the loans and 8.51% of the loans where FICO is below 740

14 of 50

Bivariate Analysis

Categorical Variables with Non Payments over 180 days

Loan Purpose

Occupancy status

Property type

15 of 50

Bivariate Analysis

Numerical Variables with Non Payments over 180 days

Fico

Original Loan-to-Value

Debt-to-Income

16 of 50

Bivariate Analysis

Numerical Variables with Non Payments over 180 days

Upb

Original rt_c

17 of 50

Bivariate Analysis

Categorical Variables with Prepayment

Loan Purpose

Occupancy Status

Property Type

18 of 50

Bivariate Analysis

Numerical Variables with Prepayment

Original Loan-to-Value

Debt-to-Income

19 of 50

Bivariate Analysis

Numerical Variables with Default 180

Upb

Original rt_c

20 of 50

Correlation Analysis

With Prepayment

With Default 180

21 of 50

Correlation Analysis

Correlation between key variables

Correlation analysis is important in the sense that we will be able to choose which variables to include in our model for survival analysis based on its results.

22 of 50

Summary

FICO: Based on our common sense, FICO is a very significant index since it highly reflects the default probability of a borrower. When FICO is high, the default probability is low. When FICO is high, the prepayment rate is high.

After the bivariate analysis, it echos our common sense. FICO and default rate have monotonic negative relationship while it has a monotonic positive relationship with prepayment rate.

​

O-LTV: OLTV should have a positive relationship with default rate. We could see in our analysis that this directly proportional relationship holds.

​

DTI: DTI should also have a positive relationship with default rate. We could also see in our analysis that this directly proportional relationship holds as well.

​

Loan Purpose: We know based on our intuition that cash out finance has the highest risk compared with purchase. Borrow to finance has the lowest risk. In our analysis, it does not align with our intuition, the cash out finance has the lowest default rate. The reason may be linking to the limited sample size.

​

​

​

​

​

​

​

23 of 50

Summary

Property Type: we assume that single family house is easy to sell. Namely, it has the lowest default risk, while the condo is riskier. The results counter our intuition. The single family has the highest default rate.

​

Occupancy Status: based on our intuition, the house for investment shall has the highest risk while Primary residence is the safest one. Based on our analysis, the results also conflict our intuition. The primary residence possess the highest risk in default rate.

24 of 50

Model Building

Week

3

25 of 50

Variable Selection

Loop Calibration: In order to test variables against their effects on KS score, we constructed a loop that measures the significance of each variable. This was done in order to allow us to get a closer look at the relationship between variables and their effects on the fit of our model.

26 of 50

Default Model

​

 

​

  • Variable Selection:

​

Intuition: intuitively using variables having economic and social impact on mortgage survival model, such as FICO, DTI, CLTV etc.

​

Spline: adding spline nodes for some variables to increase the prediction accuracy.

​

Looping: using looping to pinpoint variables highly correlated with dependent variable.

​

​

​

​

Summary:

  • KS D score: 52.
  • Average Error Rate: 10%

​

27 of 50

Default Model

28 of 50

Prepayment Model

  • Variable Selection:

​

Intuition: using variables intuitively impact the prepayment, such as FICO, original rate, relative unemployment, and the current rate in order to improve the fit of our model.

​

Looping: using looping to pinpoint variables highly correlated with dependent variable.

​

​

Summary

  • KS D score: 0.338
  • Average Error Rate: 5.84%

​

​

​

​

​

​

​

​

29 of 50

Prepayment Model

30 of 50

Summary

  • Continuously picking and combining variables to try to beat the Champion model via KS and average error rates.

​

  • Tweaking spring nodes

​

  • Trying different dummy variables

​

  • Analyzing and trying to ameliorate the result combining with other model estimating methods, such as ROC and AUC etc.

​

​

​

31 of 50

Model Validation

Week

4

32 of 50

Default Challenger 1

Our first challenger model used the variables listed on the left. We saw a high D score of .4655, but our error rates were too high.

33 of 50

Default Challenger 2

Rounds for Improvement

​

  • Reduced variables to avoid overfitting;
  • Reduced spline nodes;

​

Results:

​

  • Increased D score
  • Errors are still high

​

​

​

​

​

​

​

34 of 50

Default Challenger 3

Achieved the best model:

​

  • Further Reduced the variables with collinearity;

​

  • Improve the spline nodes

​

Result:

​

  • Garner the best model as of now.

​

​

​

​

​

​

​

​

​

​

​

​

​

35 of 50

Default Champion

36 of 50

Default Champion

37 of 50

Prepayment Challenger 1

​

For this first model, we simply used FICO scores to predict prepayment of loans. Surprisingly enough, we saw a high D score of .556, but we knew that we needed to add in more variables in order to have a more logical model.

38 of 50

Prepayment Challenger 2

​

This challenger model included more variables then the prior, such as relative unemployment, the original rate, and the current rate. We saw a D score of .338, and an error rate of .0584. While this was a better result, we still wanted to improve in both our D score and lower our error rate.

39 of 50

Prepayment Champion

Our champion model for Prepayment included variables such as FICO, DTI, CLTV, prepay rate incentive, and the relative unemployment rate. This led us to finding a D score of .344. While this is not as high as some of our challenger models, we knew that this was the most reasonable model constructed. We realized an average error rate sitting around 4.36%.

40 of 50

Prepayment Champion

41 of 50

Stochastic

Analysis�

Week

5

42 of 50

HPA Model- Parameter Trend

Fit Sigma

Series Sim

​

Sigma Sim

43 of 50

Interest Model–1 year Parameter Trend

44 of 50

Interest Model–10 year Parameter Trend

45 of 50

Summary

  • Used the FHA methodology to estimate the HPA and mortgage rate model.
  • Time horizon: 10 years for 40 quarters.
  • Matrics are generated as followers for two models:

​

​

46 of 50

Loss Distribution�

Week

6

47 of 50

Loss Distribution

6 Tranches from Fannie Mae;

Loss subordination:

(Class 1A-H:2%; Class 1M-1: 1.6%;

Class 1M-2: 1.25% ;Class IB-1: 0.7%;

Class 1B-2: 0.25% ;Class 1B-3H:0% )

Loss rate distribution:

​

  • The loss rate with the highest frequency is about 1%.
  • The far tail part has a loss rate about 2.25%;

48 of 50

Loss Distribution

Loss Rate Distribution & Tranches Loss Subordination

49 of 50

Valuation and Tranches Pricing

�

Week

7

50 of 50

CRT Tranche Price and Yields

  • The senior tranche has the lowest yield about the risk free rate-3.372%

​

  • The junior tranche has the highest yield since its price is the lowest and it undertakes the highest loss risk;

​

  • Payoff frequency is used
  • The Ideal Pricing and Yields Result: