What AI Found, What We Already Knew
Interpreting a Paper Autonomously Generated from Data by an AI-Scientist
Opening: We Handed the Data to a Machine
We ran 36 models. Across 5 methodological paradigms, drawing on 16 theories, we asked the same question of 282 cases three times over. We organized the results into two reports (a strategy edition and an interpretation edition).
Then we gave the same data to an AI-Scientist. “Analyze it yourself, and write a paper.”
The AI chose K-means clustering on its own. It autonomously carried out 5 experiments. It wrote an 8-page paper. The cost was $8, the time 20 minutes.
What that AI found overlaps exactly with one piece of the conclusion we had already reached.
Chapter 1. What the AI Saw
The Choice of Clustering
The AI had no theoretical guidance whatsoever. It knew nothing of TOE, of P-E Fit, of absorptive capacity. All we gave it were two seed_ideas (nonlinear extension of TOE, dose-response separation) as hints.
The AI generated a third idea on its own: “Classify firms with K-means, and analyze each cluster separately.” Its scores were Interestingness 9, Feasibility 9, Novelty 8. Among the four ideas, it gave this one the highest rating of all—to itself.
This is the same intuition we had pursued with T2 (LPA) in our first-round analysis. The sense that there are hidden types within the data. The suspicion that beneath the average lies heterogeneity.
Three Profiles
The 3 clusters the AI classified:
ClusternDT awarenessSmart systemsFirm size (log)Q3_difference0 (low readiness)1221.8416.12.511.5251 (high readiness)513.2425.43.341.4982 (medium readiness)701.8921.33.731.247
Here the first odd point appears. The Q3_difference of the high-readiness firms (cluster 1) is almost identical to that of the low-readiness firms (cluster 0). 1.498 vs 1.525. DT awareness is twice as high and smart systems are 60% greater, yet the training effect is the same.
This is exactly what we found in T2 (LPA) of the first-round analysis: “A more prepared firm does not necessarily learn better.” It is the very pattern the interpretation report explained as the “large-firm paradox” and the “ceiling effect.”
The Dramatic Gap of R² .002 → .336
The AI’s most impactful finding:
Cluster 0 (low readiness, n=122): R² = .002. No variable explains the training effect.
Cluster 1 (high readiness, n=51): R² = .336. Smart systems (+) and firm size (-) are strong predictors.
Cluster 2 (medium readiness, n=70): R² = .018. Again, nothing explains anything.
What was R² = .029 across the entire dataset jumps to .336 once only the high-readiness firms are isolated. 12 times higher. The AI called this “dramatic heterogeneity.”
Chapter 2. What We Already Knew
When the AI’s findings are set against our 36 analyses, the puzzle pieces fall into place.
Comparison 1: “Predictors Work Only in High Readiness”
The AI’s finding: R² is meaningful only in cluster 1.
Our findings: - T2 (LPA): The difference in training effect across profiles was not significant, but the mechanism differed by profile. - T3/T20 (QCA): Equifinality. There are multiple paths to a high training effect. ~DT_AWARE (low awareness) is also a valid path. - combo1 (RSM): The P-E Fit interaction term (p=.002) is significant only when “awareness and infrastructure are both high at the same time.”
What the AI found with K-means is another expression of what we had reached through QCA’s equifinality and RSM’s interaction term. “It works only in high readiness” and “the awareness × infrastructure interaction term is significant” are different angles on the same phenomenon.
Comparison 2: “The Negative Effect of Firm Size”
The AI’s finding: in cluster 1, firm size β = -0.839 (p<.001). Even among high-readiness firms, the larger the size, the lower the training effect.
Our findings: - T14 (ANOVA): The large-firm paradox. An adverse effect in firms of large size. - T15: SF + multiple-participation reverse synergy. The limits of additional input in firms that already have a lot. - RSM (B1): An inverted-U for firm size (p=.029), optimal at ~13 people.
The AI’s β = -0.839 corresponds to the right-hand declining segment of the inverted-U. Because high-readiness (cluster 1) firms have a large average size (log 3.34 ≈ 28 people), they have already passed the optimum of the inverted-U (~13 people). They are in the region of diminishing returns.
This is one side of what the interpretation report explained as “the dual mechanism of Liability of Smallness + ceiling effect”—and the AI captured it.
Comparison 3: “A Difference in Slopes, Not Intercepts”
The AI’s most refined finding: adding cluster dummy variables yields non-significance (p=.856), but adding cluster × variable interaction terms yields significance (p=.043).
This means that while the mean training effect of the three groups is similar (no difference in intercepts), the mechanism that determines the training effect differs (a difference in slopes).
The integrative proposition of the interpretation report was precisely this:
“The effect of SME digital transformation training is not a matter of ‘present or absent,’ but appears ‘when awareness and infrastructure meet.’”
Not “present or absent” (difference in intercepts) → consistent with the AI’s non-significant dummy (p=.856). “When they meet” (conditional) → consistent with the AI’s significant interaction term (p=.043).
Chapter 3. What the AI Missed
The reason the AI’s paper received a Reject (an internal review—it reviews what it itself wrote, based on a model built from CS-side papers) is not simply a journal mismatch. The things essential to a social-science paper are missing.
Missing 1: Theory
The AI’s paper has no theory. There is a description—“there is heterogeneity”—but no explanation of “why heterogeneity arises.”
We drew on 16 theories. P-E Fit explained the interaction term, diminishing returns the inverted-U, equifinality the multiple paths, experiential learning the dose-response. A finding without theory is no more than a pattern. That the AI found R² .002 → .336 is impressive, but without a theoretical answer to “why do predictors work only in high readiness?” it remains an unfinished finding.
The answer: it is P-E Fit. Only in firms where the fit between awareness (P) and infrastructure (E) is high does the mechanism activate by which additional organizational characteristics (smart systems, firm size) predict the training effect. In firms with low fit, the mechanism itself is “switched off.”
Missing 2: Triangulation
The AI reached its conclusion with K-means alone. We applied 5 methodologies to the same question: - LPA (person-centered) → profile types - QCA (set-theoretic) → condition combinations - RSM (nonlinear) → interaction terms / inverted-U - DID/PSM (causal inference) → treatment effects - CLPM (longitudinal) → causal direction
A single K-means silhouette of .293 may be called “reasonable,” but is that pattern confirmed by other methodologies too? Our answer was “yes, it converges across all 5.” The AI’s answer remained at the level of “I’m not sure—I changed the seed and it came out similar.”
Missing 3: The “Limits of Measurement” Insight
One of our most important findings: Training is working. It is only that Q3_difference fails to capture it.
Smart systems as DV → R² = .517
Q3_difference as DV → R² = .136
With the same IVs, simply changing the DV makes a fourfold difference in explanatory power. This is the “dose-response separation” phenomenon that converged through 5-methodology triangulation (OLS, PSM, DID, CLPM, RSM).
The AI was unable even to pose this question. It used only Q3_difference as the dependent variable and never conceived the idea, “what if we change to a different DV?” This is the absence of domain knowledge. Its ability to optimize within the data is outstanding, but the ability to imagine outside the data is not yet there.
Missing 4: Discussion
The heart of a social-science paper is the Discussion. “What does this finding add to existing theory?” “What does it mean for practitioners?” “What are the policy implications?”
The AI’s Conclusion stopped at a summary of results plus a list of limitations. The prescriptive conclusion our interpretation report reached—the word “tailored,” the argument that we must diagnose a firm’s current position and design interventions that supplement the missing dimensions—was absent from the AI.
Chapter 4. So What Does This Mean?
The Real Value of the AI-Scientist
The R² .002 → .336 that the AI found in 20 minutes and $8 is one piece of the conclusion we reached in 3 months with 36 models. About one-fifth of the whole. But the direction was exactly right.
What this means: 1. Value as an accelerator of exploration: A workflow is possible in which the AI first scans for heterogeneity through clustering, and the researcher then layers on theory and triangulates. 2. Value as a hypothesis-generation tool: The AI’s finding that “it works only in high readiness” provokes the next question, “why?” Answering that question with P-E Fit is the human’s part. 3. Value as a pattern-confirmation tool: The AI independently confirms the conclusion we reached through 36 analyses. This is a kind of methodological triangulation.
The Limits of the AI-Scientist
Absence of theory: It finds patterns but cannot explain them.
Absence of DV imagination: It optimizes only within the given DV.
Absence of literature: There is no dialogue with prior research.
Absence of Discussion: It cannot answer “so what?”
Absence of triangulation: It relies on a single methodology.
All five of these are problems of domain knowledge. Its capacity for statistical pattern recognition already approaches or exceeds the human (20 minutes vs 3 months). What it lacks is knowing “what matters in this field.”
The Next Step: AI-Scholar
Filling this gap is the development goal of the AI-Scholar.
Theory Engine: theory DB + pattern–theory matching
Literature Engine: SSCI literature search + citation generation
Analysis Engine: automated triangulation (QCA + LPA + RSM + SEM)
Writing Engine: a Theory → Hypotheses → Discussion structure
What the AI-Scientist demonstrated is possibility. In 20 minutes and $8, it can make a finding that points in the right direction. Layer theory and context onto that, and the automatic generation of a social-science paper is not impossible.
Closing: The Machine’s Eye and the Researcher’s Eye
The AI saw R² .002 → .336. We saw “when awareness and infrastructure meet.”
The AI wrote “dramatic heterogeneity.” We arrived at the word “tailored.”
The AI found the patterns. We named them, gave them meaning, and proposed what to do next.
Both of us looked at the same data, but we look at different depths. The AI’s depth is deepening fast. Yet the questions “why?” and “so what?” still belong to the human domain.
This preprint is the first social-science test of the AI-Scientist, and the starting point for developing the AI-Scholar. How the machine’s eye and the researcher’s eye can collaborate is what we must now design.
들어가며: 기계에게 데이터를 줬다
우리는 36개의 모형을 돌렸다. 5가지 방법론 패러다임으로, 16개 이론을 동원해서, 282건의 데이터에 세 번 질문했다. 그 결과를 두 개의 보고서(전략편, 해석편)로 정리했다.
그리고 같은 데이터를 AI-Scientist에게 줬다. “네가 알아서 분석하고, 논문을 써라.”
AI는 스스로 K-means 클러스터링을 선택했다. 5번의 실험을 자율 수행했다. 8페이지 논문을 작성했다. 비용은 $8, 시간은 20분.
그 AI가 발견한 것은, 우리가 이미 도달했던 결론의 한 조각과 정확히 겹친다.
1장. AI가 본 것
클러스터링이라는 선택
AI에게는 아무런 이론적 지침이 없었다. TOE도, P-E Fit도, 흡수역량도 모르는 상태. seed_ideas 두 개(TOE 비선형 확장, 용량-반응 분리)를 힌트로 줬을 뿐이다.
AI는 스스로 세 번째 아이디어를 만들었다: “K-means로 기업을 분류하고, 클러스터별로 따로 분석하자.” 점수는 Interestingness 9, Feasibility 9, Novelty 8. 네 개 아이디어 중 가장 높은 평가를 자기 스스로에게 줬다.
이것은 우리가 1차 분석에서 T2(LPA)로 했던 것과 같은 직관이다. 데이터 안에 숨겨진 유형이 있다는 감각. 평균의 이면에 이질성이 있다는 의심.
세 개의 프로파일
AI가 분류한 3개 클러스터:
클러스터nDT인식스마트시스템기업규모(log)Q3_차이0 (저준비도)1221.8416.12.511.5251 (고준비도)513.2425.43.341.4982 (중준비도)701.8921.33.731.247
여기서 첫 번째 이상한 점이 보인다. 고준비도 기업(클러스터 1)의 Q3_차이가 저준비도(클러스터 0)와 거의 같다. 1.498 vs 1.525. DT인식은 2배, 스마트시스템은 60% 더 많은데, 교육 효과는 같다.
이것은 우리가 1차 분석 T2(LPA)에서 발견한 것과 똑같다: “더 준비된 기업이 반드시 더 잘 배우지는 않는다.” 해석 보고서에서 “대기업 역설”과 “천장효과”로 설명했던 그 패턴이다.
R² .002 → .336의 극적 격차
AI의 가장 임팩트 있는 발견:
클러스터 0 (저준비도, n=122): R² = .002. 어떤 변수도 교육효과를 설명하지 못한다.
클러스터 1 (고준비도, n=51): R² = .336. 스마트시스템(+)과 기업규모(-)가 강력한 예측변수.
클러스터 2 (중준비도, n=70): R² = .018. 역시 아무것도 설명 못한다.
전체 데이터에서 R² = .029였던 것이, 고준비도 기업만 분리하면 .336으로 뛴다. 12배. AI는 이것을 “dramatic heterogeneity”라고 불렀다.
2장. 우리가 이미 알고 있던 것
AI의 발견을 우리의 36개 분석과 대조하면, 퍼즐 조각이 맞아떨어진다.
대조 1: “고준비도에서만 예측변수가 작동한다”
AI의 발견: 클러스터 1에서만 R²가 유의미하다.
우리의 발견: - T2(LPA): 프로파일 간 교육효과 차이가 유의하지 않았지만, 프로파일별 메커니즘이 달랐다. - T3/T20(QCA): 등결과성. 높은 교육효과로 가는 경로가 여러 개. ~DT_AWARE(낮은 인식)도 유효 경로. - combo1(RSM): P-E Fit 교차항(p=.002)은 “인식과 인프라가 동시에 높을 때”에만 유의.
AI가 K-means로 발견한 것은, 우리가 QCA의 등결과성과 RSM의 교차항으로 도달했던 것의 다른 표현이다. “고준비도에서만 작동한다”와 “인식×인프라 교차항이 유의하다”는 같은 현상의 다른 각도다.
대조 2: “기업규모의 부정적 효과”
AI의 발견: 클러스터 1에서 기업규모 β = -0.839 (p<.001). 고준비도 기업 중에서도 규모가 클수록 교육효과가 떨어진다.
우리의 발견: - T14(ANOVA): 대기업 역설. 규모가 큰 기업에서 역효과. - T15: SF+복수참여 역시너지. 이미 많이 가진 기업에서 추가 투입의 한계. - RSM(B1): 기업규모 역U자 (p=.029), 최적 ~13명.
AI의 β = -0.839는 역U자의 우측 하강 구간에 해당한다. 고준비도(클러스터 1) 기업은 평균 규모가 크기 때문에(log 3.34 ≈ 28명), 역U자의 최적점(~13명)을 이미 넘어선 상태. 수확체감의 영역에 있는 것이다.
해석 보고서에서 “Liability of Smallness + 천장효과의 이중 메커니즘”으로 설명했던 것의 한쪽 면을 AI가 포착한 것이다.
대조 3: “절편이 아니라 기울기의 차이”
AI의 가장 정교한 발견: 클러스터 더미 변수를 넣으면 비유의(p=.856), 하지만 클러스터×변수 교차항을 넣으면 유의(p=.043).
이것은 세 그룹의 평균 교육효과는 비슷하지만(절편 차이 없음), 교육효과를 결정하는 메커니즘이 다르다는 것(기울기 차이)을 의미한다.
해석 보고서의 통합 명제가 정확히 이것이었다:
“중소기업 디지털 전환 교육의 효과는 ‘있거나 없거나’가 아니라, ’인식과 인프라가 만날 때’ 나타난다.”
“있거나 없거나”(절편 차이) 아님 → AI의 더미 비유의(p=.856)와 일치. “만날 때”(조건부) → AI의 교차항 유의(p=.043)와 일치.
3장. AI가 놓친 것
AI의 논문이 Reject (내부리뷰 지가 쓴걸 지가 스스로 리뷰한다 CS쪽 페이퍼들로 만들어진 모델기반으로)을 받은 이유는 단순히 저널 미스매치가 아니다. 사회과학 논문에 필수적인 것들이 빠져있다.
빠진 것 1: 이론
AI의 논문에는 이론이 없다. “이질성이 있다”는 기술(description)은 있지만, “왜 이질성이 생기는가”라는 설명(explanation)이 없다.
우리는 16개 이론을 동원했다. P-E Fit이 교차항을, 수확체감이 역U자를, 등결과성이 다경로를, 경험학습이 용량-반응을 설명했다. 이론 없는 발견은 패턴에 불과하다. AI가 R² .002 → .336을 발견한 것은 인상적이지만, “왜 고준비도에서만 예측변수가 작동하는가?”에 대한 이론적 답이 없으면 그것은 미완의 발견이다.
답: P-E Fit이다. 인식(P)과 인프라(E)의 적합도가 높은 기업에서만, 추가적인 조직 특성(스마트시스템, 기업규모)이 교육효과를 예측하는 메커니즘이 활성화된다. 적합도가 낮은 기업에서는 메커니즘 자체가 “꺼져 있다.”
빠진 것 2: 삼각검증
AI는 K-means 하나로 결론을 냈다. 우리는 같은 질문에 5가지 방법론을 적용했다: - LPA (인원중심) → 프로파일 유형 - QCA (집합론) → 조건 조합 - RSM (비선형) → 교차항/역U자 - DID/PSM (인과추론) → 처치효과 - CLPM (종단) → 인과 방향
K-means 하나의 실루엣 .293은 “합리적”이라고 할 수 있지만, 그 패턴이 다른 방법론으로도 확인되는가? 우리의 답은 “예, 5가지 모두에서 수렴한다”였다. AI의 답은 “모르겠다, 시드를 바꿔봤더니 비슷하다” 수준에 머물렀다.
빠진 것 3: “측정의 한계” 통찰
우리의 가장 중요한 발견 중 하나: 교육은 작동하고 있다. 다만 Q3_차이가 그것을 포착하지 못할 뿐이다.
스마트시스템 DV → R² = .517
Q3_차이 DV → R² = .136
같은 IV인데 DV만 바꾸면 설명력이 4배 차이. 이것이 5-방법론 삼각검증(OLS, PSM, DID, CLPM, RSM)으로 수렴한 “용량-반응 분리” 현상이다.
AI는 이 질문 자체를 하지 못했다. Q3_차이만 종속변수로 사용했고, “혹시 다른 DV로 바꾸면 어떨까?”라는 발상이 없었다. 이것은 도메인 지식의 부재다. 데이터를 안에서 최적화하는 능력은 뛰어나지만, 데이터 밖을 상상하는 능력은 아직 없다.
빠진 것 4: Discussion
사회과학 논문의 핵심은 Discussion이다. “이 발견이 기존 이론에 무엇을 추가하는가?” “실무자에게 무엇을 의미하는가?” “정책적 함의는?”
AI의 Conclusion은 결과 요약 + 한계 나열에 그쳤다. 우리의 해석 보고서가 도달한 “맞춤형이라는 단어” — 기업의 현재 위치를 진단하고 부족한 차원을 보완하는 개입을 설계해야 한다는 처방적 결론은 AI에게 없었다.
4장. 그래서 이것은 무엇을 의미하는가?
AI-Scientist의 진짜 가치
AI가 20분과 $8로 발견한 R² .002 → .336은, 우리가 3개월과 36개 모형으로 도달한 결론의 한 조각이다. 전체의 약 1/5. 하지만 방향은 정확했다.
이것이 의미하는 것: 1. 탐색의 가속기로서의 가치: AI가 먼저 클러스터링으로 이질성을 스캔하고, 연구자가 이론을 입히고 삼각검증하는 워크플로우가 가능하다. 2. 가설 생성 도구로서의 가치: AI의 “고준비도에서만 작동한다”는 발견은, “왜?”라는 다음 질문을 유발한다. 그 질문에 P-E Fit으로 답하는 것은 인간의 몫이다. 3. 패턴 확인 도구로서의 가치: 우리가 36개 분석으로 도달한 결론을, AI가 독립적으로 확인해준다. 이것은 일종의 방법론적 삼각검증이다.
AI-Scientist의 한계
이론 부재: 패턴은 찾되, 설명하지 못한다.
DV 상상력 부재: 주어진 DV 안에서만 최적화한다.
문헌 부재: 선행 연구와의 대화가 없다.
Discussion 부재: “so what?”에 답하지 못한다.
삼각검증 부재: 단일 방법론에 의존한다.
이 다섯 가지는 모두 도메인 지식의 문제다. 통계적 패턴 인식 능력은 이미 인간에 근접하거나 초과한다(20분 vs 3개월). 부족한 것은 “이 분야에서 무엇이 중요한가”를 아는 것이다.
다음 단계: AI-Scholar
이 Gap을 메우는 것이 AI-Scholar의 개발 목표다.
Theory Engine: 이론 DB + 패턴-이론 매칭
Literature Engine: SSCI 문헌 검색 + 인용 생성
Analysis Engine: 삼각검증 자동화 (QCA + LPA + RSM + SEM)
Writing Engine: Theory → Hypotheses → Discussion 구조
AI-Scientist가 보여준 것은 가능성이다. 20분과 $8로 방향이 맞는 발견을 할 수 있다. 여기에 이론과 맥락을 입히면, 사회과학 논문의 자동 생성이 불가능하지 않다.
마치며: 기계의 눈과 연구자의 눈
AI는 R² .002 → .336을 봤다. 우리는 “인식과 인프라가 만날 때”를 봤다.
AI는 “dramatic heterogeneity”라고 썼다. 우리는 “맞춤형”이라는 단어에 도달했다.
AI는 패턴을 발견했다. 우리는 그 패턴에 이름을 붙이고, 의미를 부여하고, 다음 행동을 제안했다.
둘 다 같은 데이터를 봤지만, 보는 깊이가 다르다. AI의 깊이는 빠르게 깊어지고 있다. 하지만 “왜?”와 “그래서?”라는 질문은 아직 인간의 영역이다.
이 프리프린트는 AI-Scientist의 첫 번째 사회과학 테스트이자, AI-Scholar 개발의 출발점이다. 기계의 눈과 연구자의 눈이 협력하는 방법을, 이제부터 설계해야 한다.
Chae, Chungil. 2026. “What AI Found, What We Already Knew.” April 3, 2026. https://chadchae.github.io/posts_en/2026-04-03-what-ai-found-what-we-already-knew/what-ai-found-what-we-already-knew.html.


