Showing posts with label text to video. Show all posts
Showing posts with label text to video. Show all posts

Thursday, April 25, 2024

Five Top Tech Takeaways: AI Rapping Mona Lisa, Calls for AI Oversight, AI in the Banks, AI Art Receives Limited Copyright and Meta Competes with ChatGPT


Lights, Camera, Robots!


Movie Maker's Dream or Deep Fake Nightmare? Microsoft's AI Animates Art with Caution

Microsoft's latest AI innovation, detailed by their researchers last week, introduces an AI model capable of animating still images into realistic videos synchronized with audio. The VASA-1 technology can produce lifelike animations from photos, artwork, or cartoons, featuring accurate lip-syncing and naturalistic movements of the face and head. Demonstrated with a video of the Mona Lisa rapping a comedy piece by Anne Hathaway, the technology aims to enhance educational tools, assist individuals with communication challenges, and possibly create virtual companions. 

However, concerns about potential misuse for impersonation and misinformation persist. Microsoft has decided to withhold the public release of VASA-1 to ensure responsible use, aligning with practices of cautious distribution similar to their partner OpenAI's approach with the AI video tool, Sora.

Author's note: As discussed in these two posts (here and here), the possibility of creating a Hollywood studio in one's garage is becoming more realistic. Many will highlight the challenges posed by deepfakes, which are undoubtedly significant. However, there is also a positive aspect to consider. These tools could potentially enable artists to tell their stories through a cost-effective model. In fact, Tyler Perry halted his $800 million studio construction plan, influenced by the advanced visual effects achievable through OpenAI's Sora, demonstrating the shifting economic landscape of Hollywood production.

Key Takeaways:

  • Microsoft's new AI, VASA-1, animates still images into realistic videos with natural movements and synchronized audio.
  • The technology showcases potential uses in education and accessibility, yet also raises significant concerns about misuse for creating false representations.
  • Microsoft is withholding VASA-1's public release, focusing on responsible and regulated technology deployment.
(Source: CTV News)

For the original Microsoft post and more videos, see here

Former OpenAI Board Member Calls for Regulatory Oversight in AI

Former OpenAI board member Helen Toner advocated for increased transparency and regulation in the AI industry during a TED talk in Vancouver. Toner emphasized the necessity for AI companies to publicly disclose details about their technologies' capabilities and risks, and to implement robust data collection systems to address incidents. She proposed the establishment of "AI auditors" to ensure that companies are held accountable, rather than self-regulating, reflecting on her experiences and challenges, including her controversial tenure on OpenAI's board.

Key Takeaways:

  • Helen Toner, ex-OpenAI board member, stressed the importance of AI companies being transparent about their technologies and the associated risks.
  • Toner proposed the creation of independent "AI auditors" to oversee company practices and enhance accountability in the industry.
  • Reflecting on her own experiences, Toner highlighted the need for effective incident reporting mechanisms within AI companies, akin to those in aviation.
(Source: Bloomberg)

Jamie Dimon Outlines AI’s Role in JPMorgan’s Future in Shareholder Letter

In JPMorgan Chase's latest shareholder letter, the critical role of artificial intelligence (AI) in the firm's growth and operations was highlighted. Over the past decade, the firm has significantly expanded its AI capabilities, now boasting over 2,000 AI and machine learning (ML) specialists and a robust portfolio of over 400 AI-driven use cases across various business sectors like marketing, fraud, and risk management. The firm is also exploring generative AI's potential to enhance software engineering, customer service, and general productivity. Recognizing AI's importance, a new executive role—Chief Data & Analytics Officer—has been established to ensure AI and data are integral to decision-making processes company-wide.

Key Takeaways:

  • JPMorgan Chase has developed extensive AI and ML capabilities, with over 2,000 experts and 400 active AI use cases driving business improvements.
  • The firm is actively exploring generative AI applications to reimagine business workflows and enhance overall productivity.
  • A new executive role, Chief Data & Analytics Officer, has been created to integrate AI deeply into the company’s strategic and operational decisions.
(Source: JPMorgan Chase)

GenAI, Art, & Copyrights: USCO Grants Limited Copyright for AI-Assisted Work

Elisa Shupe successfully obtained copyright registration for her AI-assisted novel, "AI Machinations: Tangled Webs and Typed Words," from the US Copyright Office (USCO). Shupe extensively used OpenAI's ChatGPT while writing the book and initially faced rejection from the USCO. However, with the help of the Brooklyn Law Incubator and Policy Clinic, Shupe appealed the decision, arguing that she used ChatGPT as an assistive technology due to her disabilities. The USCO granted Shupe copyright for the selection, coordination, and arrangement of the AI-generated text, but not for the actual sentences and paragraphs. This decision is seen as a significant marker in how the USCO is grappling with the concept of authorship in the age of AI.

Key Takeaways:
  • The US Copyright Office granted Elisa Shupe copyright registration for her AI-assisted novel, recognizing her as the author of the selection, coordination, and arrangement of the AI-generated text.
  • Shupe's case highlights the nuances and challenges the USCO faces in determining the scope of protection for works produced using AI.
  • The decision to grant Shupe a limited copyright registration is seen as a compromise, as she believes she should be able to copyright the actual text of the book due to her extensive involvement in the creative process.
(Source: Wired)

Explore Meta's Latest AI Innovation: Llama 3

Meta has introduced Meta Llama 3, their newest large language model (LLM), marking a significant advancement in AI capabilities. Llama 3, which includes models with 8B and 70B parameters, boasts state-of-the-art performance in various AI benchmarks and supports a wide range of applications with improved reasoning and coding abilities. These models will soon be available across major cloud platforms and feature enhancements in trust and safety with tools like Llama Guard 2 and CyberSec Eval 2. Meta's commitment to open-source development continues with the release of Llama 3, aimed at fostering innovation and responsible use in the AI community. 

Note: Meta.ai is available in Canada and you can try it out, here: https://www.meta.ai/

Key Takeaways:
  • Meta Llama 3 introduces enhanced large language models with 8B and 70B parameters, setting new standards for AI performance and capabilities.
  • The models are part of Meta's open-source initiative, ensuring broad accessibility and encouraging community-based innovation and development.
  • Meta emphasizes responsible AI development, incorporating advanced safety features and guidelines to support secure and ethical usage.
(Source: Meta)

Author: Malik Datardina, CPA, CA, CISA. Malik works at Auvenir as a GRC Strategist who is working to transform the engagement experience for accounting firms and their clients. The opinions expressed here do not necessarily represent UWCISA, UW, Auvenir (or its affiliates), CPA Canada or anyone else. This post was written with the assistance of an AI language model. The model provided suggestions and completions to help me write, but the final content and opinions are my own.


Tuesday, April 16, 2024

Five Top Tech Takeaways: Canada's $2.4B AI bet, Adobe Goes Open, Training Data Shortage, Cdn SMBs Go Big on AI and Turnitin's Take on AI & Plagiarism

Canada Invests $2.4 Billion in AI


$2.4 Billion Infusion: Canada's Move to Spearhead AI Innovation and Safety

Canada is advancing its position in the global AI sector, as detailed by the Canadian government's announcement of a $2.4 billion investment package from Budget 2024 aimed at enhancing Canada's AI capabilities. This investment is intended to catalyze job growth, improve productivity, and ensure responsible development and use of AI technologies across various industries. The funds are allocated towards enhancing computing capabilities, boosting AI startups, supporting small to medium businesses with AI adoption, and establishing new institutes and programs for AI safety and workforce transition. These efforts underscore the Canadian government's commitment to maintaining Canada's leadership in AI innovation and providing high-quality job opportunities in the sector.

Key Takeaways:
  • The Canadian government has announced a $2.4 billion investment to strengthen the nation's AI sector, aimed at boosting job creation and productivity.
  • Investments include significant funds for computing infrastructure, support for AI startups, and programs to aid businesses and workers in adopting AI technologies.
  • The establishment of a new Canadian AI Safety Institute and the strengthening of AI legislation highlight Canada's focus on the responsible and secure advancement of AI technology.
(Source: PM Canada)

Adobe Opts For Open: Embracing OpenAI's Tools in Premiere Pro

Adobe is exploring a partnership with OpenAI and other companies as it integrates third-party generative AI tools into its Premiere Pro video editing software. This initiative aims to enhance the software's capabilities by allowing adding AI-generated objects or removing distractions with minimal manual effort. Adobe is leveraging its proprietary AI model, Firefly while considering how to incorporate external AI technologies like OpenAI's Sora. Despite the ongoing development and lack of a set release timeline, Adobe's strategy reflects its efforts to innovate amidst a competitive landscape and a significant drop in stock value this year.

Comment:  Adobe's strategic decision to make Premiere Pro open to third-party AI video makers has enabled it to avoid the pitfalls that Apple initially faced with its closed ecosystem approach to the Macintosh. Adobe has "future proofed"Premiere Pro by allowing access to third-party AI video makers. This approach contrasts sharply with Apple's early strategy with the Mac and nearly repeated with the iPhone, which restricted third-party access, limiting system functionality and user choice. By embracing openness, Adobe has enhanced its offering to video creators who want to leverage AI-generated content. 

Here, Igor Pogany walks us through the demo that Adobe has released:

Key Takeaways:

  • Adobe is integrating third-party AI tools into its Premiere Pro software, potentially enhancing video editing capabilities. This includes OpenAI, Runway ML, and PikaLabs. 
  • The company continues to use its AI model, Firefly while exploring collaborations with OpenAI and other AI developers.
  • Despite the potential of these AI tools, Adobe faces market pressures, with its stock declining by about 20% this year.
(Source: Reuters)

Turnitin Tackles AI: Insights from 200 Million Paper Reviews

In the past year, over 22 million student papers suspected of utilizing generative AI were submitted for review, according to the latest data from Turnitin, a prominent plagiarism detection company. This development follows the integration of an AI writing detection tool by Turnitin, designed to identify AI-generated content within student work. Despite the challenges of distinguishing AI-authored content from human writing, the tool has evaluated over 200 million papers, flagging 11% as containing significant AI-generated content. This surge in AI use among students underscores the evolving landscape of academic integrity and the need for sophisticated detection tools that balance effectiveness with fairness, particularly in avoiding bias against non-native English speakers.

Key Takeaways:

  • Turnitin's AI detection tool has reviewed over 200 million papers, identifying a notable percentage with significant AI-generated content.
  • The tool's development highlights the growing concern over academic integrity in the era of AI, prompting the need for reliable detection methods.
  • Issues of bias and the complexity of AI detection in academic settings remain significant, influencing institutions like Montclair State University to reassess their use of such technologies.
(Source: Wired)

AI Adoption Soars Among Canadian SMBs: A Look at the Numbers

A recent report by Float reveals a significant increase in artificial intelligence adoption among Canada's small to medium-sized businesses (SMBs), with 32% now subscribing to ChatGPT, up from just 14% a year earlier. This surge reflects a broader trend of integrating AI to enhance efficiency and productivity across various sectors, not only in mundane tasks but throughout entire organizations. According to Rob Khazzam, CEO of Float, this growth is not just a technological shift but a necessary evolution to extend operational budgets further. Despite general economic caution, with most companies maintaining flat spending levels, advertising expenses have notably doubled, indicating a readiness for growth. The report, which analyzed credit card transactions across 1,000 companies, also highlights a robust increase in spending among larger firms, signaling potential economic rebound.

Key Takeaways:
  • AI adoption among Canadian SMBs has more than doubled in a year, with 32% now using ChatGPT.
  • Businesses are applying AI broadly across functions, aiming to maximize efficiency and extend financial resources.
  • Despite cautious spending in general areas, advertising expenditures have doubled, suggesting a move towards aggressive growth strategies.
(Source: BNN Bloomberg)

The Data Dilemma: AI Giants Grapple with Training Material Shortages

OpenAI has developed its Whisper audio transcription model to transcribe over a million hours of YouTube videos for training its GPT-4 model, as reported by The New York Times. Despite legal ambiguities, OpenAI pursued this method under the belief it constituted fair use. The company is exploring the creation of synthetic data to diversify its training resources further. Meanwhile, Google and Meta are also navigating the constraints of training data availability, with Google adjusting policies to expand permissible data use and Meta considering acquisitions to secure more content. These strategies highlight the intense demand for high-quality data as AI companies strive to enhance their models' capabilities amidst growing legal and ethical scrutiny.

Key Takeaways:
  • OpenAI utilized a large volume of YouTube video transcripts, believing it to be fair use, to train its GPT-4 model.
  • The AI industry faces a critical shortage of high-quality training data, pushing companies like Google and Meta to seek creative solutions.
  • Legal and ethical challenges continue to complicate the sourcing of training data for AI models.
(Source: The Verge)

Author: Malik Datardina, CPA, CA, CISA. Malik works at Auvenir as a GRC Strategist who is working to transform the engagement experience for accounting firms and their clients. The opinions expressed here do not necessarily represent UWCISA, UW, Auvenir (or its affiliates), CPA Canada or anyone else. This post was written with the assistance of an AI language model. The model provided suggestions and completions to help me write, but the final content and opinions are my own.

Thursday, July 13, 2023

Beyond the Writer's Strike: Will AI Lead to a Renaissance of Artist Driven Content?


In our previous post, we looked at how AI has the potential to upend the way Hollywood works. With generative AI, the writers are rightfully scared about how the technology can potentially curtail their value in film production. With generative AI, I could generate a story in 35 minutes. It needed much work. However, the AI that I used was not trained on scripts. Neither was I. Imagine we both were. What stories could we generate then? 

Though AI is taking center stage in the kerfuffle, the friction has also exposed a hidden tension underlying the mass movie industry. It is the tenuous relationship between artistic expression and the commercial nature of the television and film industry. The studios that drive Hollywood only cares about recurring profits. They could not care less about art. They’ve always wanted a formula. Prompt script here. Press play on the production process. Put money in the bank and watch stock prices go to the moon. Everything else is irrelevant. Generative AI will give the studios what they want. But it will be a hollow victory. Why? Generative AI will eventually upend the Hollywood Hit Machine  as well. But before we get there, we need to discuss how Optimus Prime got into our heads.

Tapping into Pester Power: Transformers and the Deregulatory Reagan Era
Working on some side projects, I had the fortune of coming across Drawn to Television: American animated sf series of the 1980s by Lincoln Geraghty. The article explores the cartoon era of the 1980s, dominated by cartoons like Transformers, GI Joe, My Little Pony, Thundercats, and more. Geraghty ties the genesis of this genre to Star Wars. Toy companies aimed to replicate the triumph of Kenner's Star Wars action figures by creating a market through TV shows. These shows served as prolonged advertisements for an assortment of toys.

But why did this development wait until the 1980s? The Reagan Administration deregulated television and allowed toy companies to sell directly to kids of all ages and sizes. Before that, the FCC prevented such commercial interests from tapping into children's pester power.

What does this have to do with art and Hollywood?

Film critics did not think much of Transformers and the like. They saw it as “…little more than poorly drawn, glorified half-hour commercials for action figures and video games.” David Wise, a critical writer in the original Transformer series, gives us a better idea of how commercial it was. He explains that the "Rebirth" episodes were initially slated as a five-part mini-series. They were designed to introduce 92 new characters to sell as toys. He was then asked to condense the five-part story into just three episodes. Wise calculated that a new character must be introduced every 12.5 seconds. To make the storyline workable, Wise introduced groups of characters simultaneously, revealing their names and moving on – illustrating that Wise had to sacrifice the story for sales.

The Death of Optimus Prime: Killing off the Old Product Line for the New
Perhaps, the fundamental contradiction between art and commerce can be seen in the toy company’s decision to kill off Optimus Prime in the full-length movie, Transformers: The Movie, released in theatres in 1986. Wise revealed that Hasbro was disappointed with the sales of the toy-truck-robot figurine. The decision was summarized as follows:

“It was a toy show. We just thought we were killing off the old product line to replace it with new products.”

According to this cold hard logic, Optimus Prime seems to be the ultimate unscrupulous used car salesman. However, instead of peddling to adults, he sells to kids. Through his on-screen sacrifice, they could sell Rodimus Prime in his stead.


What is the reaction from young fans? According to the same consultant, Flint Dillie, who came clean about why Prime was killed off, explains how traumatic this was for children who loved the series. Kids were crying in the theatres. Families were so upset that they left during the movie. They even took to their pens, pencils, and typewriters to register their protest with the company. Hasbro gladly gave in. They had us exactly where they wanted us. This turn of events would give sagging Optimus Prime sales the needed boost.

There can't be the best way to entertain children. Specifically, it’s hard to convince a child, parent, or anyone that such an extractive relationship is healthy. How does a parent calm down a despondent child who just saw their hero killed off? It’s probably not to offer them the latest “Prime” that Hasbro has to offer. The larger point, however, is that commercially-driven content clashes not just with artistic expression but how a transaction approach to content is non-optimal for us as a whole.  

The Hollywood Hit Machine: Losing its Luster in the Age of Authenticity
The writer's strike has a limited impact on the content I usually consume, published by YouTubers, podcasts, and other enthusiasts. This shift in popular preference speaks for itself. People prefer to hear stories from real, relatable people instead of the formulaic commercial narratives churned out by the Hollywood Hit Machine.

A good proxy of the shift is the decline in cable television.

As reported by Adweek for June 2023, during prime time, FOX garners the largest average viewership with 1.49 million viewers, followed by MSNBC with 1.32 million, and CNN with 635,000 viewers. These networks collectively attract approximately 3.46 million viewers, representing about 1% of the United States estimated population of 330 million.

Another piece of evidence is the sudden and swift demise of Quibi.

Despite the company raising $2 billion, retaining the A-list of Hollywood talent, and being led by the former Disney executive Jeffrey Katzenberg, the company had shut its doors after six months. The business model was to offer short-form content in the 10 to 15-minute range – short enough to be consumed on a train ride to work. Was the pandemic, as the company claimed, the reason for its demise?

The pandemic proved to be a boom to Netflix and other streaming companies, so that's not the likely cause. Instead, what likely caused the company to crash and burn was that user-generated content was a much better source for short-form content.

Generative AI: The Great User Generative Content Amplifier?
Now we finally get to AI!

As I argue in this Medium post, generative AI is about amplification, not abdication. The post speaks to the issue of abdication from a professional perspective. A lawyer, consultant or CPA can't rely on public-facing generative AI models to do their work. Instead, it can amplify their effort by putting polish on the rough notes they have gathered.

Similarly, it’s abdication to get generative AI to produce a fictional novel in 35 minutes, hoping it will receive rave reviews. According to the Wall Street Journal, a surge in AI-generated story submissions, influenced by online videos promoting ChatGPT, led to the temporary closure of online submissions at Clarkesworld, a science-fiction magazine. Publishers, including Clarkesworld's Neil Clarke, expressed their tendency to reject these AI-written submissions, characterized by grammatically perfect but incoherent and formulaic narratives.

Using generative AI to create these types of submissions signifies an instance of “author abdication." It's the generative AI version of spam. And like we have spam filters and other "internal controls" (like the infamous proof of work concept invented to fight spam), Clarkesworld and others will need to develop similar controls to separate the good from the bad.

Instead, budding authors must work hard to conceive storylines that resonate. It could take weeks and months to sort out plot lines and characters. And you will still need to know Da Vinci Resolve, Premiere Pro, or another video editing tool.

In terms of the maturity of the tools, they have yet to arrive. However, we can see that day is quickly coming. Consider the following that is already out there.

AI Image Generation is Amazing: The current ability to generate images from a few sentences is simply the stuff of science fiction. Using stability.ai, I used this prompt “Snowy winter wonderland with a lone cabin in the distance, surrounded by frosty trees and fresh snowfall, peaceful, serene, detailed, winter landscape” to generate the following image:



AI Image Generators Enable Panning and Zoom: As explained in this video, Midjourney can generate AI images and now can allow the panning of an image. It also allows a zoom-out feature.

Professional Narrators for the Price of a Latte: In Kevin's heroic struggle, I got a professional-sounding voice to narrate the story. The cost? Eleven Labs sells this for the bargain price of $5 a month. The next tier is only $20/month.

Text to Video is Already Here: Matt Wolfe, who follows the generative AI space, compiled this video that looks at the current state of what’s out there with text to video. Lot’s to be desired. However, we’re only nine months into the Generative AI boom. The footage includes Runway ML, featured on Vox’s Recode podcast. The interview discusses how AI eliminates the need for manual labour for rotoscoping. The technique was used in the movie Everywhere All at Once, saving the production team "several hours."

The first nonsuccessful film not produced by Hollywood is still years away. However, with the rapid pace at which these tools will improve, it takes little imagination to see that the cheque is in the mail.

Reel to Real: Is There Life Beyond Hollywood?
We do not have to go far to see the types of stories that people will produce that are not driven entirely by commercial interests. Consider the historical drama DiriliÅŸ: ErtuÄŸrul. The series chronicles the rise of ErtuÄŸrul, whose son, Osman I, would establish the Ottoman State in present-day Turkey. And there are documentaries like Ava DuVernay's The 13th. The Netflix documentary explores the mass incarceration of African Americans in the US. The popularity of the ErtuÄŸrul illustrates that there is no need to make up heroes when they already exist. At the same time, the success of the 13th proves that people are interested in reality – not just fiction.

To be sure, we can expect Hollywood to continue for the foreseeable future. Cable television still attracts millions, albeit with a much-reduced viewership from its glory days. However, the shift in audience preference towards content from relatable individuals, coupled with the rise of sophisticated AI tools, indicates that the dawn of a new era in filmmaking is at hand. It's potentially a future where anyone can tell a story, where unique voices are heard, and commercial interests don't kill off characters that kids love. This technological revolution might enable a broadening of storytelling, creating space for a multiplicity of voices and narratives beyond the confines of Hollywood.

Author: Malik Datardina, CPA, CA, CISA. Malik works at Auvenir as a GRC Strategist that is working to transform the engagement experience for accounting firms and their clients. The opinions expressed here do not necessarily represent UWCISA, UW, Auvenir (or its affiliates), CPA Canada or anyone else. This post was written with the assistance of an AI language model. The model provided suggestions and completions to help me write, but the final content and opinions are my own.