What are the copyright and training-data concerns with generative AI?
Two big questions. First, inputs: training on copyrighted text, images, or code without a license raises whether that use is lawful — this is genuinely unsettled and litigated (e.g., cases involving authors, artists, and code). Some argue fair use / text-and-data-mining exceptions apply; rights holders disagree, and outcomes differ by jurisdiction. Second, outputs: models can memorize and reproduce protected content or produce near-duplicates, and there are open questions about who owns AI-generated work (the US Copyright Office has said purely AI-generated output isn't copyrightable). Practical risk management: track data provenance and licenses, prefer licensed or public-domain data, offer opt-outs, filter for regurgitation, and get legal review. As of 2026 the law is still developing, so treat this as evolving and verify current rulings.