Sunday, November 13, 2011

Final Research Report

Lime cat sez use your brain. 


What do you need to create for the final research report?   The description of the assignment is  here:
Read these instructions:
https://docs.google.com/document/d/1kQQHDh6uhIuaoWl_LWCDaAYobMGSRriBFiFUT8YYQG8/edit
USE THIS DATA: http://db.tt/JjeomqVk

Create your document and share it with Ted Now:
https://docs.google.com/spreadsheet/viewform?formkey=dC00N0padE03NTZGM3FROWlfc2ZRbkE6MQ

See newly added examples on using excel here: 
https://docs.google.com/document/d/1bsW0w9XxYwpiHOxDz-5b6c9CNPdXeHK0BOOg5_L8Ibw/edit


But here are the instructions in a nutshell:
  1. Use the data we collected to describe patterns in the variables we collected. 
    1. use your brain to find interesting things to look at
    2. Here is the link to the final data set.   please download it and save a version of your own on your computer, or to your USB drive, or to Dropbox-- wherever you prefer to store your important files: USE THIS DATA: http://db.tt/JjeomqVk
  2. look for patterns in how some variables are related to others. 
    1. Use your brain to come up with interesting variables
    2. Display the results in interesting ways
  3. Then right up what you found in a summary with the same general format as the abstract assignment. 
    1. The details of what to include are in the link above
  4. The dataset is NOW ready.  
    1. The first step is to create some basic descriptive statistics, and create some new variables that might be interesting. 
    2. Next step would be to make some graphs that would be potentially informative.
  5. This is an exploratory study--  you need to do some exploring.  Not all of your explorations will end up being noteworthy.  
    1. I will start adding examples of things you can create in excel.   I will start with how to make transformed versions of your variables.
  6. The phrase "use your brain"  is a joking way to emphasize that there are many hard parts of doing research-- but one of the most difficult is figuring out how to learn something interesting from the data. This takes effort, and is hard to teach, but is demonstrated in the research that we read for HW2.

Saturday, November 12, 2011

Creating new variables: Tips and techniques


See the google doc for info on creating variables.  Leave comments there if you have questions.

https://docs.google.com/document/d/1bsW0w9XxYwpiHOxDz-5b6c9CNPdXeHK0BOOg5_L8Ibw/edit

Here is a link to the comment data sheet as an excel that I used in the above example. Use it to practice making variables.  I will post a different data set that combines both of our data sheets once we have all the data collected. http://db.tt/VWMbFBRk

Extra Credit


Go to the Login ID variables google doc.   Sorry I was so slow getting this posted-- but here is a pretty easy chance for some extra credit.  We need to enter karma and the other login ID variables for the rest of the posters in the first 500 messages that we collected content variables for.

So-- here is the deal.   You get 10 pts extra credit for a set of 20, each person is limited to 2 sets max.


  1. Go to the google doc
  2. claim a set of 20 login ids to code (put your name in the coder field for those 20
  3. open the reddit thread
  4. search for the login id
  5. find it and click the link on the name
  6. copy their info into the spreadsheet
  7. Finish your 20
  8. take a screen shot, and add it to your homework #4 page.


Tuesday, November 1, 2011

Home Work #4



Turn in Homework #4 here:


  1. Step one:   create your google doc of your HW4. 
  2. Share with anyone with the link 
  3. Copy the link that the dialog box creates
  4. Open the following survey, and enter your name and the paste the link. 
  5. https://docs.google.com/spreadsheet/viewform?formkey=dFBzekozaS1iVTVYaWtRMV9XWGxacWc6MQ
  6. Thanks!
Homework #4 involves data collection (and analysis) from a thread in an online community called reddit.   The thread is one of the most active in recent weeks where people discussed the situation of Scott Olsen.  Scott is a two-tour veteran of the Iraq war who was hit in the head by a tear-gas canister during an 'occupy' protest in Oakland California.

Here are the instructions on HW4-- you can get started on the first two tasks as soon as you want.
https://docs.google.com/document/d/1VVpqOjxgNSFaVBfPQQ5_EIC0kTaNNrDyRqiC7Nim8LE/edit

Here is a link to the thread:






Here is the link to the OlsenMessageContentCoding
https://docs.google.com/spreadsheet/ccc?key=0ArdK55LJYgKQdHJfN3FBUk11M2ZyTklRbnh3R0JLeGc


Initial tasks:

  1. Read some of the thread.  There are over 2000 posts in the discussion so you should read a selection.  
    1. Notice that you can change the order of how posts are displayed by choosing different sort options (the default is best).   
  2. Begin identifying attributes of the discussion that we might want to measure.  
    1. attributes of the content
      1. what people talked about, media, violence, etc.
    2. attributes of how people talked to others
      1. expressed thanks, criticized another poster, 
    3. whatever else seems important
    4. Note that the posts are voted up or down by other readers and low scoring contributions are, but default, not shown.
  3. What do you think is important or interesting about this topic?
  4. What do you think is interesting about this conversation?
  5. Finally-- you might want to get a sense of related threads or conversations on the site.
  6. Finally, finally, here is a quantcast plot of the popularity of the general reddit site over the recent years. 


All of the stories on reddit are contributed by readers / users of the sites, and the links and each persons comment are all subject to voting.   

Last quiz: Hooray!


The last quiz is on lecture and reading material from week 8 (Chapters 9 Experiments + 10 Surveys) and week 9 (Chapters 11 Existing Data + 12 Quantitative Analysis).

The questions will be posted here by Tuesday evening.  The format will be similar to the previous quiz that included 3 or 4 MC questions per chapter and a couple of short answer questions.  To make sure everyone has plenty of time to complete the quiz it will not be due until class time on Thursday.

I still need to create the form for entering the quiz answers-- but here are the questions.


1.   A true experiment is defined by the following conditions:
A.   random stratified sample;  an intervention;  and a control group.
B.   random assignment to experimental condition; an intervention where you change the value of the key variable; and pre- and post tests.
C. random stratified sample; an intervention where you change the value of the key variable; and pre- and post tests.
D.  random stratified sample; an intervention where you change the value of the key variable; and a control group.

2.  Why do researchers do experiments?
A.  To thwart Perry the platypus.
B.  To measure the diversity of opinions in a population.
C.  To distinguish between multiple causal mechanisms.
D.  To test the influence of a key causal mechanism.

3.  Threats to internal validity include
A.  How events unrelated to the treatment occur during the experiment and influence the dependent variable.
B.  The maturation effect.
C.  When some research participants drop out.
D.  When the pretest measure itself affects the result.
E.  Population generalization.
F.  The placebo effect.
G.  A, B, C, D and F
H.  OMG LOL 2-Many choices
I.   A, B, C, D, E and F

4.  Why don’t sociologists use experiments more often?

5.  What is the most widely used data gathering technique in the social sciences?
A.  Experiment
B.  Use of non-reactive data
C.  Survey
D.  Field experiment

6.  A double negative is not different than the opposite of
A.  An untrue answer
B.  Wut?
C.  This question is bad because it confuses the respondent by using multiple negative constructions.
D.  Pinocchio.


7.  One disadvantage of using an open ended question is that

A.  They permit an unlimited number of possible answers.
B.  They are easier and quicker for respondents to answer.
C.  They can suggest ideas that the respondents would not consider.
D.  Responses may be irrelevant and buried in useless detail.


8.  What types of things make people less likely to cooperate with a survey?

9.  What is the key difference between using data from surveys or experiments compared to archival data, like arrest records?
A.  Experiments and surveys are designed by researchers.
B.  People think about the fact they are will be studied when they get arrested.
C. People do not think about the fact that they will be studied when they participate in a survey.
D.  When people get arrested they are not thinking about being the subject of sociological research.

10.  Non-reactive measures that emphasize the fact that that people being studied are not aware of it because the measures do not intrude.
A.  Double blind experiments
B.  Floaters
C. Unobtrusive measures
D. Erosion measures

11.  The study on SES and social capital discussed in Tuesday, Nov 1st’s lecture is an example of how
A.  Survey research can be combined with data from unobtrusive measures.
B.  Collaboration within a single department.
C.  Content analysis and social network analysis.
D.  Unobtrusive data can extend the reach of experimental studies.  

12.  What are some ways that data from online communities provide new insight into social interaction?

13.  What does it mean when the mean is larger than the mode?
A.   The distribution is normal
B.  The distribution is right skewed  (more cases with small values)
C.  The distribution is left skewed  (more cases with large values)

14.  The ability to adjust for the influence of multiple control variables is the advantage of
A.  Correlation
B.  Suppressor variable patterns
C.  Multiple regression analysis
D.  Bi-variate correlations

15.  When a relationships is described as significant it means that
A.  a relationship of that strength is unlikely to be found due to pure chance.
B.  the probability of finding a relationship of that strength is greater than one.
C.  Inferential statistics rely on principles of probability sampling.

16.   Assume that we collect a bunch of variables from the reddit discussion thread data.  What might the descriptive statistics tell us?


click here to enter your answers:
https://docs.google.com/spreadsheet/viewform?formkey=dE05dkNpQzRMTkM1eXNMYkY1cWRpZkE6MQ