Showing posts with label Viktor Mayer-Schönberger. Show all posts
Showing posts with label Viktor Mayer-Schönberger. Show all posts

Friday, November 3, 2017

Big Data Auditing Revisited: Context is King

It has been a few years since I wrote up on Big Data and the Audit.  It was one of the more popular posts with over a 1,000 hits to date.

The post looks at Big Data: A Revolution That Will Transform How We Live, Work, and Think by Kenneth Cukier and Viktor Mayer-Schönberger. I enjoyed the book as it really broke down the business impact of big data without getting in technical details of the underlying technology.

Why take a second look at big data auditing?

Big data and the accompanying analytical models are key a precursor to artificial intelligence. Machine learning algorithms that power the AI bots requires the users to analyse the problem and teach the underlying algorithm.

Part 1: Context is King

To make things a bit more digestible, I thought it would be good to divide the post into two parts. The first post is more palatable as I want to explore the second use case in a bit more detail and its relevance to today.  The second post will be a bit more controversial as I will take a look at the difficulty of applying fraud or cancer-fighting algorithms in the realm of (external) financial audit.

But let's look at the first issue: how can big data analytics give us better context? 

In the original post, I spoke discussed the use case used in Cukier and Mayer-Schönberger's work around Inrix. The book gives the example of how an investment firm is using traffic analysis, from Inrix, to determine the sales that a retailer will make and then buy or sell the stock of the retailer on that information. In a sense, the investment is using vehicular traffic as a proxy for sales. In an audit context, auditors can develop expectations of what sales should be based on the number of vehicles going around stores. For example, if sales are going up, but the number of vehicles are going down then the auditor would need to take a closer look.

What I realized from this example is that what big data can give auditors better context around things and assess reasonability of things. That is as more sensor data and other data are available to auditors to integrate into statistical models, the more they will be able to spot anomalies. 

One of the issues with Barry Minkow's ZZZBest accounting fraud was the lack of context. For more on the fraud, check this video:

I actually studied this case in my auditing class at the University of Waterloo. One of the lessons we were take away from this case was that the auditors didn't know how much a site restoration would cost on average (see the first bullet in this text on page 129). But how would an auditor be able to access such data? Even with the advent of the internet, it is not simply a matter of Googling for the information.

More recently, an accounting professor was found to have generated data fraudulently. The way he got caught was that a statistic he used didn't correspond to reality. Specifically:

"misrepresented the number of U.S.-based offices it had: not 150, as the paper maintained (and as a reader had noticed might be on the high side, triggering an inquiry from the journal)" [Emphasis added]

Again, the reader had the context to understand what was presented was unreasonable causing the study to unravel and exposing the academic fraud perpetrated by Hunton. 

What will it take to make this a reality? 

What's missing is a data aggregation tool that can connect to the private, third party, and public data feeds that an auditor can leverage for statistical analysis. Furthermore, for this to be useful to clients and the business community large are visualized depictions that enable the auditor to tell the story in a better way rather than handing over complex spreadsheets.

Of course for auditors to present such materials requires them to have deeper training in data wrangling, statistics and visualization tools and techniques. 

In the next post, we will revisit the first use case that I presented in the original post that explored how the New York City was better able to audit illegal conversions through the use of big data analytical techniques. Originally, I had thought this would be a good model to apply in the world of audit. However, I am revisiting this idea. 

Author: Malik Datardina, CPA, CA, CISA. Malik works at Auvenir as a GRC Strategist that is working to transform the engagement experience for accounting firms and their clients. The opinions expressed here do not necessarily represent UWCISA, UW, Auvenir (or its affiliates), CPA Canada or anyone else

Monday, July 18, 2016

Big Data and Predictive Policing: Can algorithms become racists?

Interesting article on Forbes by Thomas Davenport on Big Data. The articles discusses how various government, including Canadian Public Safety Operations Organization (CanOps), have used big data tools for "situational awareness". These systems draw on myriad sources of data to give users (e.g. law enforcement) the information they need to deal with a particular situation.

Here are a few points that I thought were worth noting:

Government is making strides in big data: We often think of Amazon, Google and other tech-giants as key users of this data. However, as the Davenport points out that the government is using this technology to assist with decision making. However, whether this is something that should be celebrated remains to be seen (see predictive policing below)

Privacy versus Value trade-off: He talks about how CanOps use of MASAS, the Multi-Agency Situational Awareness System, is limited by the filtering of sensitive information: "breadth of MASAS is noble, but it seems to limit its value. For example, as the CanOps website notes, because agencies are reticent to share sensitive information with other agencies, all the information shared was non-sensitive (i.e. not terribly useful)." It seems that this continues to be a theme that we had noted in back a couple years when discussing a similar trade-off the companies face when dealing with big data. As I noted in this post:

"privacy policies require the user to consent to a specific uses of data at the time they sign up for the service. This means future big data analytics are essentially limited by what uses the user agreed upon sign-up. However, corporations in their drive to maximize profits will ultimately make privacy policies so loose (i.e. to cover secondary uses) that the user essentially has to give up all their privacy in order to use the service."

Consequently, there still needs to be a solution as to how privacy can be respected but organizations can use the data they have collected to make better decisions.

Predictive Policing is an emerging reality: The sci-fi movie, Minority Report, paints a future where law enforcement arrests people before they commit crimes.


That future seems to be well on its.  Davenport mentions how "predictive policing" was introduced in 2014 to the NYPD.  He also mentions how much data is being collected by the police:

"It collects and analyzes data from sensors—including 9,000 closed circuit TV cameras, 500 license plate readers with over 2 billion plate reads, 600 fixed and mobile radiation and chemical sensors, and a network of ShotSpotter audio gunshot detectors covering 24 square miles—as well as 54 million 911 calls from citizens. The system also can draw from NYPD crime records, including 100 million summonses."

The idea of predictive policing was also raised in the book,  Big Data: A Revolution That Will Transform How We Live, Work, and Think, which I had explored in a multi-blog post series (click here for the first installment).

Andrew Guthrie Ferguson, Law professor UDC David A. Clarke School of Law, wrote an article on how that predictive policing is something that has not be really sorted in out in terms of legality. He notes:

"The open question is whether this big-data information combined with predictive technologies will create “predictive reasonable suspicion“ undermining Fourth Amendment protections in ways quite similar to the stop-and-frisk practices challenged in federal court.

In two law review articles I have detailed the distorting effects of predictive policing and big data on the Fourth Amendment and have come to the conclusion that insufficient attention has been given at the front end to these constitutional questions. New York has the chance now to address these issues before the adoption of the technology and should be encouraged by the same civil libertarians and ordinary citizens who challenged the stop and frisk policies."

His commentary highlights another limitation: big data predictions are biased based on how the data is collected. The stop and frisk policies he refers to disproportionately targeted minorities. Furthermore, policing is more focused on poor, black/hispanic neighbourhoods. Michelle Alexander documents in her book, The New Jim Crow, how this happens:

"Alexander explains how the criminal justice system functions as a new system of racial control by targeting black men through the “War on Drugs.” The Anti-Drug Abuse Act of 1986, for example, included far more severe punishment for distribution of crack (associated with blacks) than powder cocaine (associated with whites). Civil penalties, such as not being able to live in public housing and not being able to get student loans, have been added to the already harsh prison sentences."

Consequently, if the data by law enforcement is used to predict crime that essentially the targeting of minorities will continue to target such groups given that it is based on biased data. 

Technology often is seen to be a silver bullet for problems. However, we need to keep in mind that it is vulnerable to the human element that makes it. Given Microsoft's recent faux pas of accidentally allowing an AI avatar to become a Nazi, it is something that should actively be considered in the systems that are built to police and govern. 


Wednesday, November 4, 2015

Did WSJ go too far in exposing Apple employee home purchasing habits?

The WSJ published an article discussing the cost of houses in the Bay Area. As per the title of the article, "Apple Paychecks—One Reason for High Home Prices", the key culprit they highlight are the significant salaries that the Apple employees are allegedly paid.

The the data for the findings were based on the work done by Zillow completed "at the request of The Wall Street Journal" who "used census data to track down where workers in the census tract that is dominated by Apple’s Cupertino, Calif., headquarters live—primarily neighborhoods in the San Jose and San Francisco metropolitan areas". It's not clear if they relied on their own data to complete this analysis. As per the graph below, Zillow tied the rising house prices to iPhone sales.



To be fair, and abide by full disclosure principles, the article does also blame "[z]oning laws and regulatory red tape are key factors as well". However, would it be the WSJ if it didn't lay such a charge?

Where to begin? The article raises a lot of issues in terms of the role of publicly available data - regardless if it is only the census data, data gathered by aggregators such as Zillow or social media sites.

As I had written a couple of years ago, the article actually is the promise of social media to "return us to the village". In the village privacy was limited because people knew each other and any deeds or misdeeds made by the individual were quickly found out by the community. A good example of how social media accomplishes this was role of public in identifying the rioters involved in the post-Stanley cup "celebrations". If such a riot had happened in the village, the rioters would be have been held accountable in a similar manner.

The Zillow-WSJ effort is really along similar lines: if employees of a company or members of a particular guild were buying up houses and driving up prices in particular area; wouldn't people in the village know?

Furthermore, it actually is village business. We need to understand how we will live with one another how we are going to make the most of living together in this shared space called community, which requires an understanding of how the actions of one group within the community will impact others especially when it relates to a basic need like housing.

That being said, it opens up the issue of big data and its ramifications on privacy.  Although the above rationale translates well into issues relating to communal benefit it doesn't translate well into issues relating to how private entities can handle the information they were given for a specific purposes. This of course refers to the concept of "consent" well-established within privacy parlance.

The authors of  Big Data: A Revolution That Will Transform How We Live, Work, and Think raised this issue in there book. As I had noted in a previous post:

"The authors, however, raise a much more interesting point when discussing privacy in the era of big data. They highlight the conflict between privacy and profiting from big data. They note how the value of big data emerges from the secondary uses of big data. However, privacy policies require the user to consent to a specific use of data at the time they sign up ahead. This would prohibit companies from big data. However, corporations in their drive to maximize profits will ultimately make privacy policies so loose (i.e. to cover secondary uses) that the user essentially has to give up all their privacy in order to use the service. What the authors propose is an accountability framework. Similar to how stock issuing companies are accountable to the security regulators, the idea is that organizations would be accountable to a privacy body of sorts that reviews the use of the big data and ensures that companies are accountable for the negative consequences of the data.

For those of use that have been involved in privacy compliance, such an approach would make it real for companies to deal with the privacy issues in proactive manner. We saw how companies attitudes towards controls over financial reporting shifted from mild interest (or indifference) to active concern with the passage of Sarbanes-Oxley. In contrast, no similar fervour could be found the business landscape when addressing privacy issues. Although the solution is not obvious, the reality is that companies will make their privacy notices meaningless in order to reap the ROI from investments made in big data."












Wednesday, July 16, 2014

Privacy to be cast aside to make Big Data a reality?

This is the fourth and final instalment of a multi-part exploration of the audit, assurance, compliance and related concepts brought up in the book,  Big Data: A Revolution That Will Transform How We Live, Work, and Think (the book is also available as an audiobook and hey while I am at it, here's the link to the e-book ).  In the last two posts we explored the more tactical examples of how big data can assist auditors in executing audits resulting in a more efficient and effective audit. The book also examines the societal implications of big data. In this instalment, we look explore the privacy implications of big data.

What's are the privacy implications of Big Data?
In the past 3 instalments, we've explored the opportunities that big data affords to audit profession and society at large. In this article we look at the privacy implications raised by the book.

When we think of a totalitarian state we flash back to the regimes of world war II or the Soviet era. The book talks about how the East German Communist State invested vast amounts of resources on gathering data from its citizens in order to see who conformed with the state's ideology and who didn't. The book notes that East German secret police (the Ministerium für Staatssicherheit or "stasi") accumulated (amongst other things) 70 miles of documents. However, now big data analytics essentially enables corporations and governments to mine the digital exhaust people leave through social media, using their cell phones or logging into their email accounts and essentially eliminate the privacy people have.

Some may point to anonymization as a potential solution to the problem. However, the authors highlight how New York Times reporters were able to comb through anonymized data published by AOL to positively establish the identity of the users. This highlights that the powerful tools that have emerged from big data alter the privacy landscape. Consequently, privacy controls need to be rethought from this perspective.

The authors, however, raise a much more interesting point when discussing privacy in the era of big data. They highlight the conflict between privacy and profiting from big data. They note how the value of big data emerges from the secondary uses of big data. However, privacy policies require the user to consent to a specific uses of data at the time they sign up for the service. This means future big data analytics are essentially limited by what uses the user agreed upon sign-up. However, corporations in their drive to maximize profits will ultimately make privacy policies so loose (i.e. to cover secondary uses) that the user essentially has to give up all their privacy in order to use the service. What the authors propose is an accountability framework. Similar to how stock issuing companies are accountable to the security regulators, the idea is that organizations would be accountable to a privacy body of sorts that reviews the use of the big data and ensures that companies are accountable for the negative consequences of the data.

For those of use that have been involved in privacy compliance, such an approach would make it real for companies to deal with the privacy issues in proactive manner. We saw how companies attitudes towards controls over financial reporting shifted from mild interest (or indifference) to active concern with the passage of Sarbanes-Oxley. In contrast, no similar fervour could be found the business landscape when addressing privacy issues. Although the solution is not obvious, the reality is that companies will make their privacy notices meaningless in order to reap the ROI from investments made in big data.





Monday, June 16, 2014

Auditing the Algorithm: Is it time for AlgoTrust?

This is the third instalment of a multi-part exploration of the audit, assurance, compliance and related concepts brought up in the book,  Big Data: A Revolution That Will Transform How We Live, Work, and Think (the book is also available as an audiobook and hey while I am at it, here's the link to the e-book ).  In the last two posts we explored the more tactical examples of how big data can assist auditors in executing audits resulting in a more efficient and effective audit. The book, however, also examines the societal implications of big data. In this instalment, we look explore the role of the algorithmist.

Why do we need to audit the "secret sauce"?
When it comes to big data analytics, the decisions and conclusions the analyst will make hinges greatly on the underlying actual algorithm.  Consequently, as big data analytics become more and more part of the drivers of actions in companies and societal institutions (e.g. schools, government, non-profit organizations, etc.), the more dependent society becomes on the "secret sauce" that powers these analytics. The term "secret sauce" is quite apt because it highlights the underlying technical opaqueness that is commonplace with such things: the common person likely will not be able to understand how the big data analytic arrived at a specific conclusion. We discussed this in our previous post as the challenge of explainability, but the nuance here is that is how do you explain algorithms to external parties, such as customers, suppliers, and others.

To be sure this is not the only book  that points to the importance of the role of algorithms in society. Another example is "Automate This: How Algorithms Came to Rule Our World" by Chris Steiner, which (as you can see by the title) explains how algorithms are currently dominating our society. The book bring ups common examples the "flash crash" and the role that "algos" are playing on Wall Street in the banking sector as well as how NASA used these alogrithms to assess personality types for its flight missions. It also goes into the arts. For example, it discusses how there's an algorithm that can predict the next hit song and hit screenplay as well as how algorithms can generate classical music that impresses aficionados - until they find out it is an algorithm that generated it! The author, Chris Steiner, discusses this trend in the follow TedX talk:



So what Mayer-Schönberger and Cukier suggest is the need for a new profession which they term as "algorithmists". According to them:

"These new professionals would be experts in the areas of computer science, mathematics, and statistics; they would act as reviewers of big-data analyses and predictions. Algorithmists would take a vow of impartiality and confidentiality, much as accountants and certain other professionals do now. They would evaluate the selection of data sources, the choice of analytical and predictive tools, including algorithms and models, and the interpretation of results. In the event of a dispute, they would have access to the algorithms, statistical approaches, and datasets that produced a given decision."

The also extrapolate this thinking to an "external algorithmist": who would "act as impartial auditors to review the accuracy or validity of big-data predictions whenever the government required it, such as under court order or regulation. They also can take on big-data companies as clients, performing audits for firms that wanted expert support. And they may certify the soundness of big-data applications like anti-fraud techniques or stock-trading systems. Finally, external algorithmists are prepared to consult with government agencies on how best to use big data in the public sector.

As in medicine, law, and other occupations, we envision that this new profession regulates itself with a code of conduct. The algorithmists’ impartiality, confidentiality, competence, and professionalism is enforced by tough liability rules; if they failed to adhere to these standards, they’d be open to lawsuits. They can also be called on to serve as expert witnesses in trials, or to act as “court masters”, which are experts appointed by judges to assist them in technical matters on particularly complex cases.

Moreover, people who believe they’ve been harmed by big-data predictions—a patient rejected for surgery, an inmate denied parole, a loan applicant denied a mortgage—can look to algorithmists much as they already look to lawyers for help in understanding and appealing those decisions."

They also envision such professionals would work also work internally within companies, much the way internal auditors do today.

WebTrust for Certification Authorities: A model for AlgoTrust?
The authors bring up a good point: how would you go about auditing an algo? Although auditors lack the technical skills of algoritmists, it doesn't prevent them from auditing algorithms. The WebTrust for Certification Authorities (WebTrust for CAs) could be a model where assurance practitioners develop a standard in conjunction with algorithmists and enable audits to be performed against the standard. Why is WebTrust for CAs a model? WebTrust for CAs is a technical standard where an audit firm would "assess the adequacy and effectiveness of the controls employed by Certification Authorities (CAs)". That is, although the cryptographic key generation process is something that goes beyond the technical discipline of a regular CPA, it did not prevent the assurance firms from issuing an opinion.

So is it time for CPA Canada and the AICPA to put together a draft of "AlgoTrust"?

Maybe.

Although the commercial viability for such a service would be hard to predict, it would help at least start the discussion around of how society can achieve the outcomes Mayer-Schönberger and Cukier describe above. Furthermore, some of the ground work for such a service is already established. Fundamentally, an algorithm takes data inputs, processes it and then delivers a certain output or decision. Therefore, one aspect of such a service is to understand whether the algo has "processing integrity" (i.e. as the authors put it, to attest to the "accuracy or validity of big-data predictions"), which is something the profession established a while back through its SysTrust offering. To be sure this framework would have to be adapted. For example, algos are used to make decisions so there needs to be some thinking around how we would identify materiality in terms of  total number of "wrong" decisions as well as defining "wrong" in an objective and is auditable manner.

AlgoTrust, as a concept, illustrates not only a new area where auditors can move its assurance skill set into an emerging area but also how the profession can add thought leadership around the issue of dealing with opaqueness of algorithms - just as it did with financial statements nearly a century ago.




Thursday, June 5, 2014

Big Data Audit Analytics: Dirty data, explainability and data driven decision making

This is the second instalment of a multi-part exploration of the audit, assurance, compliance and related concepts brought up in the book,  Big Data: A Revolution That Will Transform How We Live, Work, and Think (the book is also available as an audiobook and hey while I am at it, here's the link to the e-book ). In this instalment, I explore another example of Big Data Audit analytics noted in the book and highlight the lessons learned from it. 

Con Edison and Exploding Manhole Covers
The book discussed the case of Con Edison (the public utility that provides electricity to New York City)  and its efforts to better predict, which of their manhole covers will experience "technical difficulties" from the relatively benign (e.g. smoking, heating up, etc) to the potentially deadly (where a 300 pound manhole can explode into the air and potentially harm someone). Given the potentially implications on life and limb, Con Edison needed a better audit approach, if you will, then random guessing as to which manhole cover would need maintenance to prevent such problems from occurring.

And this is where Cynthia Rudin, currently an associate professor of statistics at MIT, comes into the picture. She and her team of statisticians at Columbia University worked with Con Edison to devise a model that would predict, where the maintenance dollars should be focused.

The team developed a model with 106 (with the biggest factors being age of the manhole covers and if there were previous incidents) data predictors that ranked manhole covers in terms of which ones were most likely to have issues to those least likely. How accurate was it?  As noted in the book, the top 10% of those ranked most likely to have incidents ended up accounting for 44% of the manhole covers with potentially deadly incidents. In other words, Con Edison through big data analytics was able to better "audit" the population of manhole covers for potential safety issues.  The following video goes into some detail on what the team did:

What lessons can be drawn from this use of Big Data Analytics?
Firstly, algorithms can overcome dirty data. When Professor Rudin was putting together the data to analyse, it included data from the early days of Con Edison, i.e. as in 1880s when Thomas Edison was alive! To illustrate the book notes how there 38 different ways to enter the word "service box" into service records. This is on top of the fact that some of these records were hand written and were documented by people who didn't have a concept of a computer let alone big data analytics.

Second, although the biggest factors seem obvious in hindsight, we should be aware of such conclusions. The point is that data driven decision making is more defensible than a "gut feel", which speaks directly to the professional judgement versus statistical approach of performing audit procedures. The authors further point out that there at least 104 other variables that were contenders and their relative importance cannot be known without preforming such a rigorous analysis.  The point here is that for organizations to succeed and take analytics to the next level need to embrace culturally the concept that, where feasible, organizations should invest in the necessary leg work to obtain conclusions based on solid analysis.

Third, the authors highlight the importance of "explainability". They attribute to the world of artificial intelligence, which refers to the ability of the human user to drill deeper into the analysis generated by the model and explain to operational and executive management why a specific manhole needs to be investigated. In contrast, the authors point out that models that are complex due to the inclusion of numerous variables are difficult to explain. This is a critical point for auditors. As the auditors must be able defend why a particular transaction was chosen over another for audit, big data audit analytics needs to incorporate this concept of explainability.

Finally, it is but another example of how financial audits can benefit from such techniques, given the way non-financial "audits" are using big data techniques to audit and assess information. So internal and external auditors can highlight this (along with two examples identified in the previous post) as part of their big data audit analytics business case.







Monday, December 23, 2013

Big Data: Towards a "data driven audit"

This is the first installment of a multi-part exploration of the audit, assurance, compliance and related concepts brought up in the book,  Big Data: A Revolution That Will Transform How We Live, Work, and Think (the book is also available as an audiobook and hey while I am at it, here's the link to the e-book ). In this installment, I explore why this is a "must read" for those interested in data driven decision making, big data and information. I will also discuss some examples included in the book that make the case for "data driven audits". 

Why read this book?
This book is written by journalist, Kenneth Cukier, (who claims in this video to have to used the term "big data" before it was commonly used) and Viktor Mayer-Schönberger (an Internet and Governance professor at University of Oxford).


Given the background of the authors, it is an easy to digest book that gives the reader a good understanding of how access to large volumes of data and the use of correlations will change the way business is done and how society has a whole functions - without going into the technical detail of how big data is "crunched" at the back end.  The authors also discuss the following:

  • Why more is better: Algorithms improve by being exposed to more data - regardless of how messy it is. On the topic of size, it also comments how statistical sampling is a feature of an era when organizations could not wrap their arms around the data.
  • Consumer and business implications: The book is filled with examples that anyone can relate to, such as predicting whether the price of an airplane ticket will go up or down, as well as how Google uses search queries to predict flu outbreaks.
  • Enter "Datafication": It also distinguishes "datafication" versus "digitization", where the latter is making something into bits and bytes, whereas the former is something that can be analyzed by some sort of analytic engine. 
  • Potentially challenges and negative consequences of big data driven decisions: One of the challenges cited by the author is the "black box" nature of algorithms: how does a common person challenge the an algorithm, when it takes a rocket scientist to understand the algorithm itself? The authors also take the risk of explaining the  danger of subordinating human decision making to algorithms. For example, they note it would be problematic for governments to round up and quarantine people just because they looked up terms related to the flu. 
There are other interesting pieces, but I will bring them up over the next few blog posts. But if you want to get a full understanding, please buy the book - it's worth it!

The case for data driven audits
The book is filled with examples that illustrate the power of big data and how they impact business and society. However, there are a couple of examples that illustrate how financial audits can benefit from such techniques, given the way non-financial "audits" are using big data techniques to audit and assess information. 

Case 1: New York City and Auditing Illegal Conversions: As discussed in this excerpt of the book, Mike Flowers applied big data techniques to the problems of "illegal conversion" in New York city. As noted in the article, illegal conversions is the "the practice of cutting up a dwelling into many smaller units so that it can house as many as 10 times the number of people it was designed for. They are major fire hazards, as well as cauldrons of crime, drugs, disease, and pest infestation. A tangle of extension cords may snake across the walls; hot plates sit perilously on top of bedspreads. People packed this tightly regularly die in blazes". The data scientists working for Flowers, took the 900,000 property lots in the city and correlated "five years of fire data ranked by severity" against the following pieces of data:
  • Delinquency in paying property taxes,
  • Foreclosure proceedings,
  • Odd patterns in their usage of utilities, 
  • Non-payment of utilities,
  • Type of building,
  • Date building was built,
  • Ambulance visits,
  • Rodent complaints,
  • External brickwork.
By correlating all this and other information, they were able to improve the effectiveness of their 200 person inspection team from a "hit rate" of 13% to 70%. (Note: hit rate refers to conditions of building identified as being so bad that it warrants a "vacate order"). 

This is a pretty straightforward evidence for "data driven audits": financial auditors can identify correlations between financial data and non-financial data to determine which financial transactions need more scrutiny than others.

Not convinced?

Well, investors are already doing this. The book gives the example of how an investment firm is using traffic analysis, from Inrix, to determine the sales that a retailer will make and then buy or sell the stock of the retailer on that information. In a senses, the investment is using this as a proxy for sales. In an audit context, auditors can study the vehicular traffic around stores against the sales recorded against such stores and determine if there are issues worth investigating. 

Of course, this endeavor is not merely a matter of copying & pasting data from StatsCan and cobbling up a spreadsheet or two. It is a lot of hard work. However, this is not surprising to anyone who has been performing computer assisted audit techniques for the last decade or so. The challenge has always been in cleaning up the data and making it usable. Some of the challenges that the New York team of statisticians had, include:
  • Inconsistent data formats: The team had to bring together data sets from 19 different agencies. Each agency had a different way of describing location. Consequently, this has to be standardized so that each of the 19 data sets can be correlated to the same property.
  • Datafying expert intuition: The article describes how brickwork got added as an element to the correlation model. The data scientists on the team observed how the fire inspector could look at a building and know whether it was okay or not.
  • Understanding significance of each variable: Each variable must be assessed in its own right to avoid the problem of generalization. For example, rodent infestation is not uniform in its significance across New York city. As noted in the article, " A rat spotted on the posh Upper East Side might generate 30 calls within an hour, but it might take a battalion of rodents before residents in the Bronx felt moved to dial 311".
Although such challenges exist, there's a promise for big data to transform the way audits are done If you recall my blog post on the Oracle vs Google trial, I made a similar point: "the profession should use this opportunity to re-examine what aspects of technology should be a part of the audit practitioners skill set - given that is clear that society's attitude towards technology has clearly changed". Applying this to data driven audits, is actually less of a stretch as auditors have the skill set of dealing with data - instead of learning to code in Java. As a profession,  we need t o understand what data is out there and how this data can be correlated with financial data to make a more effective, efficient, and insightful audit.