
最近,我在谷歌(Google)搜索中输入了一个简单的问题:青少年每天看屏幕多久算过长?过去许多年,谷歌都会直接列出一串链接;这一次,它却给了我一段由人工智能(AI)生成的回答。这个AI先引用了一个数字,随后又对答案进行了补充和深化:它指出,相比单纯看时长,屏幕使用时间的质量和平衡可能更为重要;而究竟什么才算“过长”,可能还取决于青少年的睡眠、运动、课业压力和情绪状态。
于是我又搜索了另一个问题:我是否应该每天服用阿司匹林?这一次,AI给出了相关医学信息,提醒了潜在风险,并表示若我能提供年龄和病史,它还能给出更有针对性的指导。
这些回答都不错,但真正引起我兴趣的是:它们其实属于不同类型的回答。
围绕AI回答的讨论,一直主要聚焦于准确性:系统给出的回答是否准确?这一点当然重要,但准确性只是检验标准之一。面对不同类型的回答,用户需要运用不同的标准进行判断。
我发现,把AI回答归入一套“答案分类体系”会很有帮助。这套体系包含四大类,分别是:事实型、解释型、建构型和策略型。事实型回答通常可通过某个信息源加以核验;解释型回答即便准确,也仍然体现了其对核心证据的取舍;建构型回答哪怕逻辑严密,也可能并不适合特定的使用者;而一份文笔优美的策略性文件,则未必符合真实情况。然而,AI往往会用同样流畅、同样权威的形式呈现这四类回答,导致其中的差异极易被忽视。
我是弗吉尼亚大学(University of Virginia)的大学图书馆馆长。目前,我正牵头在全美范围内推动面向图书馆专业人员的AI能力建设,也广泛开展AI素养方面的咨询工作。我最早是在《学术图书馆学杂志》(Journal of Academic Librarianship)上提出了这套类型体系。
当然,这四个类别并非彼此完全隔绝的“孤岛”。AI给出的某条回答完全可能同时涵盖其中几种类型。尽管如此,我将在下文中逐一解析每类回答,并提供相应的判断指南,以帮助你决定某条回答是否可以直接采用,还是需要进一步核查。
谷歌给出的是哪类回答?
事实型回答提出的是原则上能够通过证据来核实的陈述。例如:弗吉尼亚大学建于哪一年?黄金的化学符号是什么?
要判断一条事实型回答是否足够可靠并能直接使用,应根据合适的信息源来核实。如果回答中引用了来源,应顺着链接去查证,而不是简单地把AI的回答本身当作证据。
解释型回答同样基于证据,但往往没有唯一结论。比如:青少年每天看屏幕多久算过长?远程办公能否提高生产率?答案取决于哪些证据被纳入、哪些被遗漏,以及如何理解其中的分歧。
在最初回答我的“屏幕时间”问题时,谷歌提示每天两小时是青少年的上限。但随后它又补充说,儿科学界的指导方针更看重屏幕使用的质量和具体情境,而不是简单以时长来衡量。美国儿科学会(American Academy of Pediatrics)指出,针对青少年的屏幕使用时间并没有明确的推荐标准,重点应放在屏幕使用的具体类型,以及这类活动可能挤占了哪些其他活动。一个看似只需数字回答的问题,最终却需要进行解释判断。
因此,在评估解释型回答时,不能仅仅局限于核对事实。你还需自问:系统重点强调了哪些证据?忽略了哪些证据?是否还存在另一种站得住脚的解释?此时,向搜索引擎提出一个有用的后续追问会很有帮助:“支持另一种结论的最有力证据是什么?”
建构型回答不是被“发现”的,而是被“创造”出来的。比如,让AI起草求职信、撰写悼词、设计教案,或者重新组织一段文字——这些任务都不存在唯一正确的结果。
你可以从目的、受众和口吻这三个维度来评判这类回答。一篇悼词可能在语法上无可挑剔,但读起来却完全不像致辞者本人的语气,或者无法在台下的家属心中产生共鸣,它也可能未能准确捕捉到逝者的特质。在阅读此类回答时,务必将这些实际效果考虑在内。
策略型问题探讨的是“该怎么做”。比如:我应该每天服用阿司匹林吗?我应该买下这套房子吗?这类回答将客观信息与主观判断相结合,涉及对目标、风险、权衡取舍以及个人实际情况的综合考量。
我搜索阿司匹林的经历,恰好说明了背景情境的重要性。谷歌不仅提示了潜在风险,建议我咨询医疗专业人士,还表示如果我能提供年龄、心血管病史以及出血风险等信息,它就能提供量身定制的解答。这种谨慎态度与美国预防服务工作组(U.S. Preventive Services Task Force)的指南不谋而合。该指南指出,是否开始服用低剂量阿司匹林来预防心脏病发作和中风,应当因人而异,并在心血管获益与出血风险之间仔细权衡。
面对策略型回答,你应该先问自己:系统还需要掌握哪些信息,其建议才能合理地适用于你本人?同时,还要考虑其中的利害关系、备选方案,以及是否应有专业人士介入。以阿司匹林问题为例,一个有效的追问会是:“关于我的年龄、病史或出血风险,哪些具体细节可能会改变这项建议?在作出决定前,我应该和医生讨论哪些问题?”最终的决定权依然在你手中,因为你需要亲自承担选择的后果。
得到答案后的首要问题
我一开始提出的这些问题,只是普通的谷歌搜索。我并没有打开聊天机器人,但由AI生成的回答就这样直接呈现出来,末尾还附有相关链接。
这些回答确实颇具实用性。谷歌补充了背景信息,承认了问题的复杂性,并表示若我能提供更多资料,它还能给出定制化指导。然而,即便是在同一条回答内部,答案的类型也可能发生变化。转述某项医疗指南的内容,和判断这项指南如何适用于特定个人,完全是两码事。一段行文流畅的回答能够在不同类型间自由切换,而语气却没有任何明显改变。
因此,作为用户,应尽量去辨别AI在生成一条回答时,究竟完成了何种类型的“脑力劳动”。你需要考量其解释是否具有说服力,或者其建议是否契合你的实际情况。
所以,在追问AI的答案是否正确之前,不妨先抛出一个更基础的问题:这究竟是哪一种答案?因为答案的类型,将指引你采取下一步行动。
本文作者利奥·S·罗,现任弗吉尼亚大学图书馆馆长
本文经The Conversation授权,依据知识共享许可协议转载。(财富中文网)
译者:刘进龙
审校:汪皓
最近,我在谷歌(Google)搜索中输入了一个简单的问题:青少年每天看屏幕多久算过长?过去许多年,谷歌都会直接列出一串链接;这一次,它却给了我一段由人工智能(AI)生成的回答。这个AI先引用了一个数字,随后又对答案进行了补充和深化:它指出,相比单纯看时长,屏幕使用时间的质量和平衡可能更为重要;而究竟什么才算“过长”,可能还取决于青少年的睡眠、运动、课业压力和情绪状态。
于是我又搜索了另一个问题:我是否应该每天服用阿司匹林?这一次,AI给出了相关医学信息,提醒了潜在风险,并表示若我能提供年龄和病史,它还能给出更有针对性的指导。
这些回答都不错,但真正引起我兴趣的是:它们其实属于不同类型的回答。
围绕AI回答的讨论,一直主要聚焦于准确性:系统给出的回答是否准确?这一点当然重要,但准确性只是检验标准之一。面对不同类型的回答,用户需要运用不同的标准进行判断。
我发现,把AI回答归入一套“答案分类体系”会很有帮助。这套体系包含四大类,分别是:事实型、解释型、建构型和策略型。事实型回答通常可通过某个信息源加以核验;解释型回答即便准确,也仍然体现了其对核心证据的取舍;建构型回答哪怕逻辑严密,也可能并不适合特定的使用者;而一份文笔优美的策略性文件,则未必符合真实情况。然而,AI往往会用同样流畅、同样权威的形式呈现这四类回答,导致其中的差异极易被忽视。
我是弗吉尼亚大学(University of Virginia)的大学图书馆馆长。目前,我正牵头在全美范围内推动面向图书馆专业人员的AI能力建设,也广泛开展AI素养方面的咨询工作。我最早是在《学术图书馆学杂志》(Journal of Academic Librarianship)上提出了这套类型体系。
当然,这四个类别并非彼此完全隔绝的“孤岛”。AI给出的某条回答完全可能同时涵盖其中几种类型。尽管如此,我将在下文中逐一解析每类回答,并提供相应的判断指南,以帮助你决定某条回答是否可以直接采用,还是需要进一步核查。
谷歌给出的是哪类回答?
事实型回答提出的是原则上能够通过证据来核实的陈述。例如:弗吉尼亚大学建于哪一年?黄金的化学符号是什么?
要判断一条事实型回答是否足够可靠并能直接使用,应根据合适的信息源来核实。如果回答中引用了来源,应顺着链接去查证,而不是简单地把AI的回答本身当作证据。
解释型回答同样基于证据,但往往没有唯一结论。比如:青少年每天看屏幕多久算过长?远程办公能否提高生产率?答案取决于哪些证据被纳入、哪些被遗漏,以及如何理解其中的分歧。
在最初回答我的“屏幕时间”问题时,谷歌提示每天两小时是青少年的上限。但随后它又补充说,儿科学界的指导方针更看重屏幕使用的质量和具体情境,而不是简单以时长来衡量。美国儿科学会(American Academy of Pediatrics)指出,针对青少年的屏幕使用时间并没有明确的推荐标准,重点应放在屏幕使用的具体类型,以及这类活动可能挤占了哪些其他活动。一个看似只需数字回答的问题,最终却需要进行解释判断。
因此,在评估解释型回答时,不能仅仅局限于核对事实。你还需自问:系统重点强调了哪些证据?忽略了哪些证据?是否还存在另一种站得住脚的解释?此时,向搜索引擎提出一个有用的后续追问会很有帮助:“支持另一种结论的最有力证据是什么?”
建构型回答不是被“发现”的,而是被“创造”出来的。比如,让AI起草求职信、撰写悼词、设计教案,或者重新组织一段文字——这些任务都不存在唯一正确的结果。
你可以从目的、受众和口吻这三个维度来评判这类回答。一篇悼词可能在语法上无可挑剔,但读起来却完全不像致辞者本人的语气,或者无法在台下的家属心中产生共鸣,它也可能未能准确捕捉到逝者的特质。在阅读此类回答时,务必将这些实际效果考虑在内。
策略型问题探讨的是“该怎么做”。比如:我应该每天服用阿司匹林吗?我应该买下这套房子吗?这类回答将客观信息与主观判断相结合,涉及对目标、风险、权衡取舍以及个人实际情况的综合考量。
我搜索阿司匹林的经历,恰好说明了背景情境的重要性。谷歌不仅提示了潜在风险,建议我咨询医疗专业人士,还表示如果我能提供年龄、心血管病史以及出血风险等信息,它就能提供量身定制的解答。这种谨慎态度与美国预防服务工作组(U.S. Preventive Services Task Force)的指南不谋而合。该指南指出,是否开始服用低剂量阿司匹林来预防心脏病发作和中风,应当因人而异,并在心血管获益与出血风险之间仔细权衡。
面对策略型回答,你应该先问自己:系统还需要掌握哪些信息,其建议才能合理地适用于你本人?同时,还要考虑其中的利害关系、备选方案,以及是否应有专业人士介入。以阿司匹林问题为例,一个有效的追问会是:“关于我的年龄、病史或出血风险,哪些具体细节可能会改变这项建议?在作出决定前,我应该和医生讨论哪些问题?”最终的决定权依然在你手中,因为你需要亲自承担选择的后果。
得到答案后的首要问题
我一开始提出的这些问题,只是普通的谷歌搜索。我并没有打开聊天机器人,但由AI生成的回答就这样直接呈现出来,末尾还附有相关链接。
这些回答确实颇具实用性。谷歌补充了背景信息,承认了问题的复杂性,并表示若我能提供更多资料,它还能给出定制化指导。然而,即便是在同一条回答内部,答案的类型也可能发生变化。转述某项医疗指南的内容,和判断这项指南如何适用于特定个人,完全是两码事。一段行文流畅的回答能够在不同类型间自由切换,而语气却没有任何明显改变。
因此,作为用户,应尽量去辨别AI在生成一条回答时,究竟完成了何种类型的“脑力劳动”。你需要考量其解释是否具有说服力,或者其建议是否契合你的实际情况。
所以,在追问AI的答案是否正确之前,不妨先抛出一个更基础的问题:这究竟是哪一种答案?因为答案的类型,将指引你采取下一步行动。
本文作者利奥·S·罗,现任弗吉尼亚大学图书馆馆长
本文经The Conversation授权,依据知识共享许可协议转载。(财富中文网)
译者:刘进龙
审校:汪皓
I recently typed a simple question into Google search: How much screen time is too much for teenagers? Instead of presenting links, as Google had been doing for many years, it gave me an AI-generated answer. The artificial intelligence agent cited a number, then complicated that reply, noting that quality and balance of time could matter more than the number of hours, and that “too much” time could depend on a teenager’s sleep, exercise, school demands and mood.
I tried another search: Should I take a daily aspirin? This time the AI answer presented me with medical information, warned about risks and offered more tailored guidance if I provided my age and medical history.
These were good replies. What interested me was that they were different kinds of replies.
Debate about AI answers has focused on accuracy: Did the system get the answer right? That matters, but accuracy is only one test. Each kind of answer requires a user to judge something different.
I find it useful to sort AI answers into an “answer typography” of four broad types: factual, interpretive, constructive and strategic. A factual claim can often be checked against a source. An interpretation can be accurate and still reflect choices about which evidence matters. A construction can be well reasoned and still be wrong for the person receiving it. A beautifully written strategic document may not be true. Yet AI presents all four types of answers in much the same fluent, authoritative form; the differences are easy to miss.
I’m university librarian and dean of libraries at the University of Virginia who leads national efforts to develop AI competencies for library professionals, and I consult widely on AI literacy. I first proposed the typography in the Journal of Academic Librarianship.
The four categories are not airtight boxes. A response from an AI agent can reflect several types. That said, I describe each type of answer below, and offer guidance for deciding whether a reply is ready to use or needs more investigation.
Which answer is Google giving you?
A factual answer makes a claim that can, in principle, be checked against evidence. When was the University of Virginia founded? What is the chemical symbol for gold?
To determine whether a factual answer is robust enough for you to use, verify the claim against an appropriate source. If the answer cites a source, follow that link instead of simply treating the answer itself as proof.
An interpretive reply is built on evidence, but there is not a single takeaway. How much screen time is too much for a teenager? Does remote work raise productivity? The answer depends on what evidence is included, what is left out and how disagreement is understood.
Google’s initial answer to my screen-time question indicated that two hours was a limit for teenagers. Then it noted that pediatric guidance puts more weight on the quality and context of screen use than on simple hours. The American Academy of Pediatrics says there is no exact recommended amount for teens and emphasizes the kind of screen use and what activities it might be displacing. A question that looked numerical turned out to require interpretation.
To assess interpretive answers, do more than check facts. Ask yourself what evidence the system emphasized, what it left out and whether another defensible interpretation exists. A useful follow-up question to present to the search engine is: “What is the strongest evidence for a different conclusion?”
Constructive answers are made rather than discovered. Ask AI to draft a cover letter, write a eulogy, suggest a lesson plan or reorganize a paragraph – there is no single correct result.
You can judge the response by considering purpose, audience and voice. A eulogy can be grammatically perfect and still sound nothing like the person delivering it, or it may land flat on family members hearing it. It may not capture the deceased person well, either. Consider these kinds of effects as you read.
Strategic questions ask what to do. Should I take a daily aspirin? Should I buy the house? The answers combine information with judgment about goals, risks, trade-offs and personal circumstances.
My aspirin search shows why context matters. Google warned about risks, told me to consult a medical professional and offered more tailored information if I provided my age, cardiovascular history and risk of bleeding. That caution matches the U.S. Preventive Services Task Force guidance. It says the decision to start low-dose aspirin for prevention of heart attacks and strokes should be individualized and weigh cardiovascular benefit against bleeding risk.
For strategic answers, ask what the system would need to know before its advice could reasonably apply to you individually. Consider the stakes, the alternatives and whether a qualified person should be involved. For the aspirin question, a useful follow-up would be: “What details about my age, medical history or bleeding risk could change this advice? What should I discuss with my doctor before deciding?” The final judgment remains yours because you are the person who has to live with the outcome.
The first question after an answer
My questions began as ordinary Google searches. I did not open a chatbot. The AI-generated responses simply arrived, and links were appended.
The responses were useful. Google added context, acknowledged complications and offered tailored guidance if I supplied additional information. Within each response, though, the type of answer could change. Reporting what a medical guideline says is different from deciding how it applies to a particular person. A fluent response can move between those types of answers without a noticeable change in voice.
As a user, try to recognize what kind of intellectual work the AI agent did for a response you receive. Consider whether the interpretation is persuasive or the advice fits your circumstances.
Before asking whether an AI answer is right, ask a more basic question: What kind of answer is this? The type will tell you what to do next.
Leo S. Lo, Dean, University of Virginia
This article is republished from The Conversation under a Creative Commons license. Read the original article.