Thursday, April 16, 2009

Google Scholar API and PaperCube

A Google Scholar API doesn't exist yet, but some people are asking for it. Reading a paper last night I had an idea for what would be a cool use of a Google Scholar API, which is taking all the citations in a reference section of a paper and adding in a citation count. I wonder if that could be done even without a formal API?

What would be great would be a color coding of the references list at the end of the paper that picked out the papers with the most citations ...

Interestingly, reading the full list of the google scholar API request discussion I found mention of PaperCube, a SproutCore/SVG system that runs in a browser developed by an Apple Engineer and Master's student at Santa Clara called Peter Bergstrom. This is much much cooler than my concept, and if it was running on top of a Google Scholar API would be a total killer app for academics. Being able to navigate the web of academic citations is what every academic must currently do in their head, knowing which are the key papers that have been cited in a field. At the moment, moving into a new field is hit and miss as often the papers you get back from a search by title may be of relatively little importance to the field. The existing citation counts in Google Scholar are helpful, but you still have to conduct multiple searches to build up a picture of the field; not to mention that the workings of the Google Scholar citation calculation are black box. If we are lucky Google will release the API, make the citation calculation transparent, and something like PaperCube will run on top and make all academic:s lives that little bit easier ....

An empirical investigation of knowledge creation in electronic networks of practice (Chou & Chang, 2008)

I'm reading this paper as part of my look at social psychology inspired analysis and design of online communities.

In this paper the authors are concerned with knowledge creation in electronic networks. They describe knowledge as "information combined with experience, context, interpretation and reflection", which goes a little against the situated cognition position of knowledge as process rather than entity; however the authors go on to discuss how knowledge creation involves social and collaborative processes.

The authors present a paradox that while electronic networks allow information to be shared quickly with many other individuals, and may foster knowledge creation; the lack of directives or organizational routines in networks that span organizational boundaries may hinder efficient knowledge creation. In order to better understand the knowledge creation process in electronic networks the authors develop a series of hypotheses based on the theories of social capital (Lin, 2001) and the theory of planned behaviour (TPB; Ajzen, 1991).

They test these hypotheses by creating a questionnaire on attitudes to knowledge creation (KC). The reliability of this instrument is checked by getting responses from a mixed group of industries in Taiwan. A refined questionnaire was then distributed to members of a legal professional association that made extensive use of electronic networks, and the results from this survey were analyzed and used to create a model of the relationships between the different constructs, as shown in the diagram above.

The hypotheses in the paper are about how different constructs such as reputation and centrality of an individual affect attitudes towards KC, intention to conduct KC and KC behavior itself. The majority of the author's hypotheses are born out by the results of the survey (and an analysis based on posting logs). Some hypotheses are contradicted such as those that knowledge self-efficacy and reputation positively effect attitudes towards KC. Also, network centrality appears to impede the intention to create KC.

Overall this paper includes a rigorous approach to developing a survey instrument and analyzing it in terms of specific constructs. The approach is similar to the one I read about in Ma & Agarwal (2007) and I am guessing must be a standard approach in the business/organizational sciences. This approach clearly avoids some of the standard criticisms leveled at data collected in surveys, but it does not appear to avoid the problem that the survey results rely on self-assessment. An individual might believe they have a positive attitude to knowledge creation , but the reality may be different. Some of the constructs used rely on relatively objective analysis (i.e. centrality), so it would be beneficial to have further analysis to determine the relationship between individual's perceptions of their behaviors and actual behaviors. Also I'd like to see the actual questionnaire, but presumably it was in Taiwanese, and so there would be no simple way for me to get a feel for the "reliable constructs" expressed within it.

This paper (cited by 2 according to Google Scholar at time of this post) appears to be following in the footsteps of the highly cited Wasko & Faraj (2005) and so it would probably be instructive to read that paper.

References

Ajzen, I. (1991), “The Theory of Planned Behavior,” Organizational Behavior and Human Decision Processes, 50(2), pp. 179-211.

Lin, N. (2001), Social Capital, Cambridge University Press, Cambridge, UK.

Ma, M., and Agarwal, R. (2007), “Through a Glass Darkly: Information Technology Design, Identify Verification, and Knowledge Contribution in Online Communities,” Information Systems Research, 18(1), pp. 42-67.

Wasko, M. M., and Faraj, S. (2005), “Why Should I Share? Examining Social Capital and Knowledge Contribution in Electronic Networks of Practice,” MIS Quarterly, 29(1), pp. 35-57.

Wednesday, April 15, 2009

Constructing Networks of Action-Relevant Episodes (CN-ARE): An In Situ Research Methodology (Barab et al., 2001)

Describes a way of analyzing the spread of concepts during learning where the focus is on the participation trajectories of multiple actors rather than on the minds of individual learners. Trajectories are represented as networks of activity involving material, individual and social components. To this end observations of learner interactions are broken into episodes which are coded and then represented as nodes in a network.

'Situated Cognition' theory informs this approach; the idea that knowing refers to an activity and not a thing; that it is always contextualized and not abstract. This made me ask isn't there some sort of knowledge that we like to have in abstract form, e.g. addition, but I guess even if a learner has abstracted the general concept of addition it is still contextualized in terms of the extent to which the culture in which they find themselves values mathematical skill, and the situations in which it is expected to be used; i.e. any cognition is a complex social phenomenon.

The authors present a number of methodological contexts for their own methodology.
  • Interaction Analysis - developing coding schemes to describe interactions
  • Network theory - mapping things into nodes and links
  • Activity theory - seeing things in terms of nested activity systems, i.e. participants using components to act on objects
  • Actor Network theory - tracing interactions
It seems like the nodes are little pseudo-'activity theory' elements, with each node being defined in terms of an issue at hand, an initiator, a participant, a resource and a practice. Seems to me that trying to code the practice up front could be premature; I would have thought what the practice was would emerge after subsequent analysis of a network. The diagram above shows the initiators of particular practices as circled numbers, and other participants as just numbers. Lines show relations, and shading indicating some common theme.

Nodes or Action-Relevant Episodes are delimited by a change in theme, activity, or initiator. In the example in the paper of a virtual solar system related class activity a new node might be created when the issue at hand switched from eclipses to animation, the practice from modeling to Socratic questioning, or from one student to another. A large CN-ARE diagram is presented showing the tracer network for eclipse-related nodes, but it is not clear to me what insight is necessarily derived from using this particular representation. It is clearly helpful to have the detailed coding of the learner's activities if one wishes to understand the learning process rather than just the end results, but it would be nice if there was an explicit example of how being able to highlight nodes of a particular type in CN-ARE led to particular insights. Although the authors do refer to another paper where they:
Barab, Barnett, Yamagata-Lynch, Squire, and Keating (2002) used the CN–ARE methodology to identify the frequency of occurrences related to a particular tracer, and then used Activity Theory (Engeström, 1987) to contextualize these in terms of the larger context (activity system).
So perhaps I should read that, although a quick skim of it reveals that, as they say in the above quote, that this is about frequency of occurrences, and activity systems analysis; rather than about the network diagrams. I can intuit the value of the network diagrams, but would really like to see a concrete example of an insight that they helped support, i.e. one that wouldn't have been easily discovered through another representation.

Cited by 39 at time of post according to Google scholar

Monday, April 13, 2009

The Human Infrastructure of Cyberinfrastructure (Lee et al., 2006)

This paper describes how the authors conducted an 18-month ethnographic study of a large-scale distributed biomedical cyberinfrastructure project. Based on this study the authors argue that 'human infrastructure' is shaped by traditional, as well as new, team and organizational structures. The authors suggest 'human infrastructure' as an alternative perspective to 'distributed teams' as a way of understanding distributed collaboration.

The authors use a focus on human infrastructure to describe the social conditions and activities that constitute the emergence of infrastructure. The idea is that recently there have been projects that try to promote the development of cyber-infrastucture, e.g. shared databases, IT to support multi-institutional collaboration; big efforts to create applications, software and tools to support big science. The human infrastructure consists of the programmers, researchers, developers and so forth who have experience in the difficulties of developing tools to operate in such complex environments.

This paper focuses on FBIRN (Functional Biomedical Informatics Research Network), a multi-institutional project with the goal of developing tools to make multi-site functional MRI (Magnetic Resonance Imaging) studies a common research practice. The authors describe different social groups in FBIRN, including the importance of traditional organizations (e.g. hospitals and universities), and the lack of clarity on the part of individuals as to their degree of membership in particular virtual groupings (e.g. task forces and working groups). Individuals have a restricted view of the whole project, which the authors suggest may be advantageous since the complexity of the entire project is too great for any one individual to follow. Individuals also make extensive use of personal networks and networks arising from other related projects. Some examples of the human infrastructure perspective appear to be:
  • New Practices, Old Conventions: Creation of new clinical assessment tool from multiple existing previous ones
  • Experimentation and Negotiation: Over time FBIRN developed a new process for developing experiments
  • Sharing Data: both legal and ontological issues and the challenges of keeping everyone up to date on the development of shared standards
They describe the amorphous dynamic nature of the human infrastructure, and that the formation of BIRN is recursive. However I am not sure about the use of this term. They say:
"The infrastructure is used by people to negotiate work, and in response to these interactions, the shape of the infrastructure itself is continually negotiated and changed."
and cite an example of how a statistics working group was formed out of a calibration working group, and how the individuals in the statistics working group were operating independently until a re-organization by the working group chairs. However none of this suggests a recursive structure. A dynamic and reflective one certainly, but not recursive, but perhaps I am taking the term too literally.

The authors say that they are employing 'human infrastructure' as a lens to understand the work of FBIRN; a new way to understand organizational work, in contrast to traditional organizational structures, distributed teams or networks. However, even having read the paper twice I am a little unclear about what the 'human infrastructure' perspective is supposed to be. I guess I am unfamiliar with the existing literature of traditional organizational structures. Having read Wenger's Communities of Practice I tend to think of all human endeavours in this general fashion, i.e. there being a messy cludge of multiple types of networks and contexts. The authors argue that theirs is a departure from a picture of traditional, hierarchical organizations being replacced with dynamic, networked organizational forms, in that they see the story of FBIRN making sense by combing perspectives from traditional organizations, distributed teams and personal networks.

The authors make a few recommendations such as encouraging the embracement of fluid organizational structures, that different groups will require different sort of organizational and instrumental support. The problem with any recommendations is that there does not appear to be any assessment in the paper regarding the effectiveness of the FBIRN organisation. Early on the paper we are told that the project has:
"successfully developed de novo tools for multi-site functional MRI studies, for data collection, management, sharing, and analysis"
However, we only have the word of the authors of the paper that FBIRN's achievements should be viewed in a positive light. One might ask whether the same results could not have been achieved faster and more efficiently with other types of organizational structures? Personally I would much rather work in a fluid organizations, but I wonder what basis there is for arguing that it is necessarily better.

There is mention in the paper of how some individuals found the non-traditional organization challenging in so far as they couldn't hold collaborators accountable. Although there was suggestion of some alternative techniques being developed to address this issue nothing in the paper constitutes evidence that the particular organization of FBIRN is necessarily a good one. The authors are applying an ethnomethodological approach of attending physical and virtual meetings as observers, reading emails and conducting interviews. While this clearly allows large amounts of data to be gathered, and the study is ongoing, I am not sure what I can really say with much certainty having read the paper. My personal interest would be in what technologies were being used for collaboration, e.g. teleconferencing and virtual meeting software, wikis, email lists etc. Not that I don't want to focus on the human infrastructure, but I want to understand how the humans are using the different type of technology to serve the individual goals as they relate to their traditional organizational role, their personal networks and their other project affiliations.

Perhaps I am being too hard on an unfinished study that is as yet only a conference paper, but while it was very interesting to hear about how FBIRN operates, I have not taken away a clear understanding of what it means to focus on human infrastructure as opposed to distributed teams; which I think is the key perspective shift that the paper is advocating. I guess I would need to read literature on analysis of distributed teams to achieve a better understanding.

The success of projects of such a scale would appear to be very hard to assess. How does one know that the collaborations taking place are successful? Presumably interviewing all the participants and finding that they all agree that the collaboration is successful is one approach, but I would want some additional, perhaps more objective measure. Maybe there is no such measure. However in the absence of some measure of success or failure it seems difficult to argue that one perspective is necessarily better than another.

Thinking about my own interest in design patterns for virtual organizations, this paper makes me feel that perhaps I am focusing much too small. Is the group management interface in a wiki really going to have much impact on big distributed collaborations on things like FBIRN? The big socio-technical decisions would appear to be more at the level of what mailing lists to set up, which out of the box content management system to use (e.g. wikimedia or drupal) and how frequently to have face to face meetings etc.

[N.B. This paper has been cited by 21 according to google scholar at time of this posting.]

Monday, April 6, 2009

impressed by getsatisfaction automated answer search

I find myself impressed by the automated answer search in the getsatisfaction customer feedback management system . It checks an existing database of problems so that you may be able to find the solution to your problem without having to actually post yourself. This seems like a great idea to avoid multiple posts asking the same question and repeated answers from admins or expert users saying "please refer to existing answer at blah blah blah", which I have noticed a lot on things like MySQL forums. I was considering a similar approach for the disCourse discussion tool at one point, but never got round to implementing it. It seems like a community pattern with great potential. I can't find any scholarly papers that refer to it, and otherwise my experience at getsatisfaction has not been amazing, to the extent that I haven't managed to get answers to the issues I am having related to mapufacture's geopress project. I would be very interested to hear about any similar features available in other web applications or other software; especially if there was some citable academic research on the same.

Social and Economic Incentives in Google Answers

"Social and Economic Incentives in Google Answers" is a research paper by Sheizaf Rafaeli, Daphne R. Raban and Gilad Ravid; presented at the ACM Workshop on Sustaining Community: The role and design of incentive mechanisms in online systems, Sanibel Island. As of this post it has been cited by 10 other authors according to Google Scholar; although 6 of these are self citations.

The authors describe their own analysis of pricing and behaviour in the Google Answers system and compare with the analysis of Edelman (2004). Edelman found that more specialized answerers earned less per-hour. Edelman explained that when a researcher stays within a particular field they lose opportunities in other fields. The workshop paper refers to some interesting economic properties of information such as:
Information is expensive to produce and cheap to reproduce (Bates, 1989; Shapiro
and Varian, 1999).
and
... (information's) value is revealed only after consumption (Shapiro and Varian, 1999; Van Alstyne, 1999).
They also note that:
Behavioral research revealed that the value of information is derived from perceptions of at least three central elements: cost, quality, and ownership (Toften and Olsen, 2004; Raban and Rafaeli, 2005).
The researchers' initial inspection of the data suggested a correlation between economic incentive and amount of questions answered; while tips were only very weakly correlated, and the socially constructed ratings were not correlated at all.

Of course I should have been reading and citing the author's subsequent journal paper: Rafaeli, S., Raban, D., Ravid, G. How Social Motivation Enhances Economic Activity and Incentives in the Google Answers Knowledge Sharing Market. Int. Journal of Knowledge and Learning 3, 1 (2007), 1--11, but I had the earlier conference paper printed out ... but then again the technical paper ends rather abruptly with the conclusion that:
... when interaction is present, the social parameters of rating and comments contribute incentives to the formation of participation, beyond the role of economic incentives ...
Which I didn't quite follow from the conference paper, but this gets explained in more detail in the subsequent journal paper. It seems that if one considers only those questions that generated a discussion (at least one comment) then tips and social ratings are correlated to the question being answered. More detailed analysis in the journal paper of the mean ratio of comment per answer indicated that comments from the community given before answers from experts is correlated with the likelihood of answers from experts, but only at the level of individual experts. This is in contrasts with a reduced number of questions answered by experts where comments are present, if the analysis is not at the expert level, but across the entire site. The authors explain this at the site level as follows:
... if sufficient help was provided by a comment, there is less need or room for an answer. Ethical experts will not post a paid answer where an informative comment was submitted.
whereas the converse relationship at the individual expert level in this fashion:
Experts seem to be drawn to questions that generate much interest, comments, activity and tend to answer those questions more often. This may fill a social need but may also serve an economic purpose of enhancing the expert's reputation by getting exposure to more eyeballs.
This contradiction at the different levels of analysis only makes sense to me when one takes into account that comments may, or may not, answer the question posed. If the answer is contained in the comments they will serve as a dis-incentive to expert answers, whereas comments that do not answer the question are indicative of interest in the answer and as such serve as an incentive.

The authors consider two possible explanations to this and the related finding of increased levels of tipping for individual experts when questions have many comments:
... comments may enhance the overall perceived quality of all knowledge provided, as answer or comments, so the asker becomes more inclined to tip ... (or) ...it may be that the asker feels some social pressure by the presence of the comment contributors which leads him/her to provide a tip as a social norm.
It seems to me that the latter explanation is the more likely, and the authors go on to cite a number of papers on the subject of social facilitation, which suggests they also lean towards a similar explanation. However it is difficult to imagine an objective experiment or analysis which would tease these two possible explanations apart. Although one might experiment with auto-generated comments or something similar.

This makes me think that critical mass on an individual piece of information or across an entire site may be a function of the likelihood with which individuals perceive their activity as being observed and/or approved of my others. For example, if lots of others, particularly existing colleagues of respected individuals, are undertaking certain activities in regard to a post, or towards a site, then this increases the likelihood of an individual participating, and when this likelihood (combined with frequency of encounter) reaches a tipping point over the system as a whole then the critical mass needed to achieve a thriving online community will be achieved.

I originally found this paper because it cited Ling et al. (2005), which I consider a seminal study on motivation in online communities. I am particularly interested in how certain design patterns influence participation in online communities, and Ling et al was one of the first papers that I found that took an experimental approach to this subject; something I perceive to be missing from the existing research on online community design patterns.

The journal paper has a couple of non self citations at the time of this post http://scholar.google.com/scholar?cites=13346321855162398665

There do appear to have been a few other analyses of the Google Answers since this paper was published http://scholar.google.com/scholar?q=%22Google+Answers%22

Interestingly google answers shut down on December 1st 2006. Other services such as Mahalo Answers, and UClue have sprung up in its place. There were also some non-monetary services that ran in parallel with Google Answers and continue to thrive such as Yahoo Answers and Answer Bag. Wikipedia and this blog post have a good discussion of some of the reasons for the shutdown.

Thursday, April 2, 2009

Collaborating with Web Applications

There are many web applications available that try to support online collaboration. I have been involved for several years in building and maintaining one system that was originally in PHP and was migrated to Rails, and is currently stuck at Rails 1.0, while the Rails world has moved on, most recently to Rails 2.3.

Over time I have become aware of many other systems such as Plone, Drupal and hosted services like basecamp and huddle. I often draw a parallel as follows:

PHP -> Zend -> Drupal
Ruby -> Rails -> Radiant
Python -> Zope -> Plone

Where at the base we have a programming language, next we have a framework that makes it easy to produce web applications in that particular language, and then finally we have systems that run as complete content management system (CMS) out of the box, support some basic funcationality like login, wikis etc., and also support the addition of extra functionality through some plugin framework.

Of course the picture is not as simple as this model in that firstly I don't think that Drupal runs on Zend (although there is a zend module for Drupal), and there are many frameworks for each language and many CMSs for each language and framework. However the named elements in the above schema are ones that have become close to the defacto standard in their bracket, and it is instructive to think about the pros and cons of setting up a CMS or collaboration system using each.

Writing something from scratch in PHP, Ruby, Python or other language is really a hell of a lot of work and you will be maintaining it yourself, and it will be difficult to get open source community traction given the other existing frameworks written in those languages, so really the options are pick a framework, or pick a CMS. Picking a CMS means less up front work and potentially getting support from the community associated with it, but the particular CMS may not match your current needs. Fortunately these systems can be extended through plugins, but if the underlying system is not a good match you'll have trouble. Working directly in the frameworks to make your own application give you a lot of flexibility but there tends to be a lot of re-inventing the wheel through making things like password reset functionality, which you will then have to maintain. One interesting possibility here is to incorporate a framework like google connect or facebook connect:

http://www.google.com/friendconnect/
http://www.elevatedrails.com/articles/2009/01/02/announcing-facebooker-support-for-facebook-connect/
or even to go for, potentially closed source, alternatives like basecamp or huddle. Interestingly huddle is now available through LinkedIn and Facebook. http://www.linkedin.com/static?key=developers_widgets

and there are interesting possibilities for using twitter
http://apiwiki.twitter.com/REST+API+Documentation

Alternatives to basecamp according to http://www.slideshare.net/carsonified/cheaponomics include
http://www.goplan.info and there is an even bigger list here: http://www.whybasecampsux.org/#alternatives

Using a hosted service leaves you at the mercy of the hosted service, and so it is tempting to set up an existing CMS, but these require maintenance too, however with plugin frameworks they do allow a degree of customization. To get more flexibility one can build in a framework like rails, but as I have found out the real challenge there is keeping things up to date, which is not as simple as it sounds. In our project we stuck with Rails 1.0 in order to keep pushing features out and avoid having to spend time adapting to a changing codebase, but we built up a lot of technical debt in the process, and maintaining our Rails 1.0 is now a big handicap. Also the lack of a pluggable framework means that involving other coders is messy, although that said I haven't really experimented with the pluggable systems enough to correctly analyze the associated pitfalls.

Anyway, so this is a rather incoherent post, but I wanted to collect together my thoughts and associated links on one page ...