Showing posts with label econometrics. Show all posts
Showing posts with label econometrics. Show all posts

Wednesday, July 24, 2013

Spurious regression, illustrated

Source:


There is a quadratic relationship between the percentage of males in a country who smoke and per capita health care spending. The relationship is statistically significant, as is shown in the table on the right.

The moral of the story? Don't conclude anything from regressions on country-level data.


The problem with analyses like this one is shown in the picture below. There is a very strong positive relationship between a country's income and the amount of money that is spent on health care: as income increases, health care spending increases. There is also a relationship - though a more complex one - between income and smoking. The poorest of the poor cannot afford cigarettes. As income increases, smoking rates increase. But in the most affluent countries, rates of smoking are lower again.

Monday, August 13, 2012

On correlated variables

* If A is correlated with B and B with C, does A need to therefore be correlated with C? 
* No. 
clear 
set obs 10000 
gen A = rnormal()  
gen C = rnormal()
gen B = A+C 
corr A B C

Tuesday, July 17, 2012

Seasonal adjustment

Some economic trends (such as unemployment and consumption of certain goods) take on a seasonal component; i.e., the trends consistently go up or down every year in a cycle. Hence, to remove such contamination and examine the true underlying trend, such trends must be reported only after executing seasonal adjustment (overview here, in-depth guide to the actual process here).

Examples of trends that need to be seasonally adjusted, from this post on the Conversable Economist:



Datagasm: World Bank links up to Stata!

From the World Bank site:
Stata is a statistical computing package widely used in the business and academic worlds. We use it at the World Bank and it’s great to see a new version of the wbopendata module that gives Stata users direct access to much of the data on data.worldbank.org... 
It’s important to have convenient access to the best data available. The wbopendata module connects to the World Bank Open Data API and provides direct access to the latest version of the Bank’s data though the Stata interface – there’s no unnecessary downloading or management of data needed.
The new version of wbopendata lets you:
  • Access 1,000 new indicators, bringing the total up from 4,200 to 5,300 time series.
  • Access the metadata of the downloaded series: including indicator definitions, the organization and/or agency responsible for its collection, and links to available supporting information.
  • Easily link the indicators downloaded to maps from within Stata.
  • Access data in three Stata-supported languages: English, Spanish or French
The wbopendata module lets you connect to information from over 256 and regions since 1960, the accessible datasets include:
To begin, just type in Stata: 
ssc install wbopendata
Further instructions can be found here. Cannot wait to try it out!

Monday, July 9, 2012

Econometrics tips of the day

1. Stata's predict command (from Econometrics by Simulation):
* Stata's predict command is an extremely useful command for many purposes.
* In this post we will go through how it works. And manually program in long hand some of the things it does.
* First let's start with OLS
* Imagine the underlying population model Y = g(x1, x2)
* Now imagine an estimator Y = f(x1, x2)
* What most estimations do is they take the Y and the xs and estimate some variant of f.
* In the linear case Y = b0 + b1x1 + b2x2 + u
* Most estimation commands attempt to estimate b0, b1, and b2. Which is great!
* But after estiamting b0, b1, and b2 what we may ask,
* "How does u look? Does it look normal, thus justifying the use of OLS?"
* We may also ask, "How does the estimated y look? This is often not particularly interesting since it is purely linear but often 'yhat' the predicted y is used in post estimation techniques."
Approximating unknown (continuously differentiable) functions by using a Taylor (MacLaurin) series expansion is common-place in econometrics. However, do you ever pause to recall that such approximations are only locally valid - that is, valid only in a neighbourhood of the (possibly vector) point about which the approximation is made? 
Unlike some other types of approximations - such as Fourier approximations - they are not globally valid.

Sunday, July 8, 2012

Probit vs. logit

Repost from Econometrics by Simulation:

* This is a comparison of how well the logit does relative to the probit when the data is generated from the assumptions underlying the probit.

* First let's generate data that is consistent with the probit assumptions

clear
set seed 101
set obs 1000

* x is the explanatory variable
gen x = rnormal()*(1/2)^(1/2)

* u is the error
gen u = rnormal()*(1/2)^(1/2)

* y is the unobservable structural y
gen y = x + u

* In order to do a probit correctly the underlying distribution has to be standard normal (which is not a restriction so long as you remember when generating values.)

* This is why rnormal*(1/2)^(1/2) -> var(y)=var(x+u)=(1/2)*1+(1/2)*1=1
sum y
* Pretty close to 1 in the sample

* y_prob is the probability of observing a 1 given
gen y_prob = normal(y)

* y is the actual binary draws
gen y_observed = rbinomial(1,y_prob)

* now let's try estimating probit first
probit y_observed x
  * Save the estimated coefficient to a local macro
  local coef_probit = _b[x]

* let's predict the probabilities
predict y_probit
  label var y_probit "Probit fitted values"

* Now let's estimate the logit
logit y_observed x
* Save the logit to a local macro
local coef_logit = _b[x]

* predict the probabilities from the logit
predict y_logit
  label var y_logit "Logit fitted values"

* We can see that both the probit and the logit are almost identical
two (line y_logit x, sort) (line y_probit x, sort)


di "It is a somewhat well known property that probits and logits are in practice almost linearly equivalent."
di "The ratio of probit to logit is: `=`coef_probit'/`coef_logit''"

reg y_probit y_logit
* Check out that R-squared!

* So what does all of this practically mean?

* Feel free to switch between probit and logit whenever you want.  The choice should not generally significantly affect your estimates.

* Note: for mathematical reasons sometimes it is easier using one over the other.

* Finally, if you want to recover the original coefficient on x the best thing to do is to take the average partial effect (APE)

probit y_observed x

di (1/2)^(1/2)

test x==.70710678
* The probit results get fairly close but we reject the null

Thursday, June 7, 2012

Featured paper of the day: Banks and securities markets

"The evolving importance of banks and securities markets" (Demirguc-Kunt, Feyen, Levine)
This paper examines the evolving importance of banks and securities markets during the process of economic development. We find that as countries develop economically, (1) the size of both banks and securities markets increases relative to the size of the economy, (2) the association between an increase in economic output and an increase in bank development becomes smaller, and (3) the association between an increase in economic output and an increase in securities market development becomes larger. The results are consistent with theories predicting that as economies develop, the services provided by securities markets become more important for economic activity, while those provided by banks become less important.
Nice paper using quantile regression techniques. (I wish I could have written this for work!) At least now I have a greater appreciation of the applications of quantile regression in finance.

Sunday, June 3, 2012

With great data come great responsibilities

A thoughtful piece on the powers and limits of data. While aimed more toward "data journalists", economists and econometricians may do well to take heed of the following advice:
Data is not a force unto itself. Data clearly does not literally create value or change in the world by itself. We talk of data changing the world metonymically – in more or less the same way that we talk of the print press changing the world. Databases do not knock on doors, make phonecalls, push for institutional reform, create new services for citizens, or educate the masses about the inner workings of the labyrinthine bureaucracies that surround us. The value that data can potentially deliver to society is to be realised by human beings who use data to do useful things. The value of these things is the result of the ingenuity, competence and (perhaps above all) hard work of human beings, not something that follows automatically from the mere presence and availability of datasets over the web in a form which permits their reuse. 
Data is not a perfect reflection of the world. Public datasets (unsurprisingly) do not give us perfect information about the world. They are representations of the world gathered, generated, selected, arranged, filtered, collated, analysed and corrected for particular purposes – purposes as diverse as public sector accounting, traffic control, weather prediction, urban planning, and policy evaluation. Data is often incomplete, imperfect, inaccurate or outdated. It is more like a shadow cast on the wall, generated by fallible human beings, refracted through layers of bureaucracy and official process. Despite this partiality and imperfection, data generated by public bodies can be the best source of information we have on a given topic and can be augmented with other data sources, documents and external expertise. Rather than taking them at face value or as gospel, datasets may often serve as an indicative springboard, a starting point or a supplementary source for understanding a topic. 
Data does not speak for itself. Sometimes items in a database will stand by themselves, and do not require additional context or documentation to help us interpret them – for example, when we consult transport timetables to find out when the next train leaves. But often data will require further research and analysis in order to make sense of it. In many ways official datasets resemble official texts: we need to learn how to read and interpret them critically, to read between the lines, to notice what is absent or omitted, to understand the gravity and implications of different figures, and so on. We should not imagine that anyone can easily understand any dataset, any more than we would think that anyone can easily read any policy document or academic article. 
Data is not power. Data may enable more people to scrutinise official activities and transactions through more detailed, data-driven reportage. In principle it might help more people participate in the formulation of more evidence based policy proposals. But the democratisation of information is different from the democratisation of power. Knowing that something is wrong or that there is a better way of doing things is not the same thing as being in a position to fix things or to affect change. For better or for worse flawless arguments and impeccable evidence are usually not sufficient in themselves to affect reform. If you want to change laws, policies or practices it usually helps to have things like implacable advocacy, influential or high profile supporters, positive press attention, hours of hard graft, bucketloads of cash and so on. Being able to see what happens in the corridors of power through public datasets does not mean you can waltz down them and move the furniture around. Open information about government is not the same as open government, participatory government or good government. 
Interpreting data is not easy. Furthermore there is a tendency to think that the widespread availability of data and data tools represent a democratisation of the analysis and interpretation of data. With the right tools and techniques, anyone can understand the contents of a dataset, right? Here it is important to distinguish between different orders of activity: while it is easier than ever before to do things with data on computers and on the web (scrape it, visualise it, publish it), this does not necessarily entail that it is easier to know what a given dataset means. Revolutionary content management systems that enable us to search and browse legal documents don't mean that it is easier for us to interpret the law. In this sense it isn't any easier to be a good data journalist than it is to be a good journalist, a good analyst, a good interpreter. Creating a good piece of data journalism or a good data-driven app is often more like an art than a science. Like photography, it involves selection, filtering, framing, composition and emphasis. It involves making sources sing and pursuing truth – and truth often doesn't come easily. Amid all of the services and widgets, libraries and plugins, talks and tutorials, there is no sure-fire technique to doing it well.

Saturday, June 2, 2012

Links worth reading

1. Alvin Roth on the future of repugnance (essay).

2. Dave Giles griping about linear probability models:
It's not the bias that worries me in this particular context - it's the inconsistency. After all, the MLE's for the Logit and Probit models are also biased in finite samples - but they're consistent. Given the sample sizes that we usually work with when modelling binary data, it's consistency and asymptotic efficiency that are of primary importance.