为什么产线声音检测这么难なぜ生産ラインの異音検査は難しいのかWhy Sound Inspection on the Line Is Hard
声音往往是装配与零件问题最早暴露的信号,但把这份判断稳定交给机器,比想象中难。音は組付けや部品の不具合が最も早く表に出る手がかりですが、その判断を機械に安定して任せるのは容易ではありません。Sound is often the earliest sign of an assembly or part problem — yet handing that judgment to a machine reliably is harder than it looks.
装配松动、零件毛刺、轴承异常,常在尺寸还合格时就先在声音里露苗头。产线想把「听」纳入检查工序,落地却接连撞上三堵墙。組付けの緩み、部品のバリ、軸受の異常は、寸法がまだ公差内のうちから音に現れます。「聞く」ことを検査工程に組み込みたいという要求は続いてきましたが、実装しようとすると三つの壁にぶつかります。Loose assemblies, burrs and bearing faults show up in the sound while every dimension is still in tolerance. Plants have long wanted listening as a formal inspection step, but three obstacles get in the way.
靠老师傅的耳朵熟練者の耳に頼るRelying on a seasoned inspector
老师傅的耳朵确实灵,但标准因人而异、随疲劳漂移。更关键的是判定不留痕迹,无法复盘也无法复制到下一条线。熟練者の耳は鋭いのですが、基準は人により異なり疲労でずれます。さらに判定の記録が残らず、振り返りも次のラインへの複製もできません。A seasoned inspector hears a great deal, but the criterion differs by person and drifts with fatigue. It also leaves no trace, so it can be neither reviewed nor copied to the next line.
靠声级与频谱的阈值规则音圧レベルやスペクトルのしきい値に頼るRelying on level and spectrum thresholds
阈值只擅长抓「变大、变响」。现场大量异常却是能量没变、音色变了:多一点摩擦的沙沙,多一点周期的咔嗒。しきい値が得意なのは「大きくなる」変化だけです。現場の異常の多くは大きさが変わらず音色だけが変わり、擦れ感や周期的なカタつきが加わります。Thresholds only catch what gets louder. Much of what goes wrong keeps its loudness and changes timbre: a faint rubbing texture, a periodic click.
靠有监督 AI 学缺陷欠陥を教師あり学習で覚えさせるTeaching an AI from labelled defects
学分类的前提是缺陷够多够全。产线恰恰相反:良品占绝大多数,缺陷稀少,等样本攒够产线早已换了一轮。分類の学習は、欠陥の数と種類がそろっていることが前提です。現場はその逆で、良品が大多数を占め欠陥は稀にしか出ず、そろう頃にはラインは別の状態になっています。Classification assumes enough defects of enough kinds. The line is the opposite: good parts dominate, defects are rare, and by the time samples suffice the line has moved on.
| 现场遇到的情况現場で起こることWhat happens on the line | 传统做法为什么难従来手法が苦手な理由Why conventional methods struggle | 对算法的真正要求アルゴリズムに求められることWhat the algorithm must do |
|---|---|---|
| 该分辨什么,本来就说不清何を手がかりにすべきか、そもそも決めきれないWhat to listen for is not obvious in the first place | 指标由人想出来,想不到就漏掉指標は人が思いつくしかなく、抜けが出るIndicators must be thought up by people, so gaps remain | 模型自己从数据里学出该分辨什么手がかりをモデルがデータから学ぶThe model learns from data what to attend to |
| 缺陷罕见且不断出新欠陥は稀で、種類も増え続けるDefects are rare, and new kinds appear | 学习先要缺陷样本学習には欠陥サンプルが前提Learning needs defect samples first | 无标注也能立判据欠陥ラベルなしで判定基準を作るBuild a criterion without labels |
| 能量不变、音色变了大きさは同じで音色が変わるSame loudness, different timbre | 总量指标分辨不出総量的な指標では捉えられないAggregate indicators cannot tell | 分辨音色与时间纹理音色と時間的な質感で聞き分けるDiscriminate on timbre and texture |
| 背景与批次持续漂移背景騒音とロット差が変化するBackground and incoming lots drift | 固定阈值逐渐失准固定しきい値は次第にずれるFixed thresholds go out of alignment | 判据随现场更新判定基準が現場に追随するThe criterion moves with the line |
| 一条线一个脾气ラインごとに癖が異なるEvery line has its own character | 一套参数难通吃一組の設定では賄えないOne setup cannot suit all | 按线挑合适的分辨方式ラインごとに適した分け方を選ぶEach line picks what suits it |
不是更多人工规则,而是会自己学的深度学习系统人手のルールを増やすのではなく、自ら学ぶ深層学習システムをNot more hand-written rules, but a deep learning system that learns for itself
在缺陷稀少、现场持续漂移的条件下仍然稳定可用。欠陥が乏しく現場が変化し続ける条件下でも、安定して使えることを目指しています。It stays usable where defect samples are scarce and the line keeps drifting.
三条设计原则三つの設計原則Three Design Principles
针对上一节的三堵墙,本引擎定下三条不退让的原则。前節の三つの壁に対して、本エンジンは三つの原則を崩さずに保っています。Against those three obstacles, the engine holds to three principles.
三条原则决定后面各章的做法:没有标注靠什么学,扩展时怎么不动已用的判定,为什么不压成一个指标。三つの原則が以降の各章の方法を決めています。サンプルなしで何を学ぶのか、能力を広げるとき運用中の判定をどう保つのか、なぜ一つの指標にまとめないのか、に答えます。The three principles decide every method described later: what the model learns from without samples, how a judgment already in use survives new capability, and why nothing is compressed into one indicator.
不靠缺陷标注,用自监督学习欠陥ラベルに頼らず、自己教師あり学習でLearn Without Defect Labels: Self-Supervised
模型自己出题、自己检验,从无标注声音里学出正常声音的表征(模型自己形成的内部表示):能量起伏、共振位置、敲击衰减。判据是学出来的,不是人写的规则。モデルが自ら課題を作り自ら確かめる形で、ラベルなしの音から正常な音の表現(モデル自身が形づくる内部表現)を学びます。エネルギーの変動、共振の位置、打撃後の減衰です。判定基準は学習で得たものであり、人が書いた規則ではありません。The model sets its own task and checks its own answer, learning a representation of normal sound — its own internal description — from unlabelled audio: how energy moves, where resonances sit, how a strike decays. The criterion is learned, not written by hand.
无需缺陷样本欠陥サンプル不要No defect samples needed预训练的深度听觉模型保持不变,只新增读出方式事前学習済みの深層聴覚モデルは動かさず、読み出し方の追加で能力を広げるKeep the Pre-Trained Deep Listening Model Fixed; Extend How It Is Read
底层模型定版后权重不再改动,接新产线只重新拟合上层的读出与门槛。不是换掉耳朵好的人,而是请他多讲一种说法。下の層のモデルは版が確定した後は重みを変更しません。新しいラインでは上の層の読み出しとしきい値を取り直すだけです。耳のよい人を替えるのではなく、別の言い方でもう一度説明してもらうということです。Once its version is fixed, the lower model's weights are never changed; a new line only means re-fitting the readouts and thresholds above it. It is not replacing the person with the good ear, but asking them to describe what they heard in another way.
底层不重训基盤は再学習しないNo base retraining用多个互补视角分辨同一段声音同じ音を複数の相補的な視点で捉えるOne Sound, Several Complementary Views
有的缺陷在音色轮廓上一下就能分辨,有的只在时间纹理里,有的要贴人耳的感受才说得清。挤进一个指标会互相稀释,各视角独立判断后合并。音色の輪郭で分かる不具合、時間的な質感にしか出ない不具合、人の耳の感じ方に沿って初めて説明できる不具合があります。一つの指標に押し込めば互いに薄め合うため、各視点が独立に判断してから統合します。Some faults are obvious in the overall timbre, some appear only in fine temporal texture, some only when the sound is read as an ear would feel it. One indicator would let them dilute each other, so each view judges independently before the judgments are merged.
互补优于压缩相補性を優先Complementary, not compressed- 既有基准继续有效,新增读出可先只记录,一致后再纳入判定。
- 扩展可随时关掉退回原状;底层没动,差异只会来自新增部分。
- 既存の基準はそのまま有効で、追加した読み出しはまず記録のみに留め、一致を確認してから判定に組み込めます。
- どの拡張も無効化して元の状態に戻せます。基礎の層は動かしていないため、差異の原因は追加部分に限られます。
- The existing baseline stays valid, and a new readout can run in record-only mode until its behaviour is confirmed.
- Any extension can be switched off to return to the previous state, and since the base layer is untouched, any difference comes from what was added.
本手册中的几个说法本書で用いる言い方A few terms used in this manual
听得清:同时用粗细不同的尺度聞き分ける:粗さの異なる複数のスケールでHearing Clearly: Several Scales at Once
一段声音要先被听清楚,才谈得上被判断。本节讲这套深度学习引擎是怎么听的。音はまず聞き分けられて初めて判断できます。本節では本エンジンがどのように聞くかを説明します。A sound must first be heard clearly before it can be judged. This section explains how the engine listens.
分辨一段声音,有个躲不开的取舍:时间上分得越准,频率上就分得越糊;反过来也一样。真正的问题不是选哪个折中点,而是为什么必须只选一个。音を分けて捉えるには避けられない取捨があります。時間を細かく分けるほど周波数は分けられなくなり、その逆もまた同じです。本当の問題は、どこで折り合うかではなく、なぜ一方しか選べないのかという点です。Resolving a sound involves an unavoidable trade-off: the finer the resolution in time, the coarser it becomes in frequency, and the other way round. The real question is not where to compromise, but why only one can be chosen.
时间分得准,频率就分不清時間を細かく分けると、周波数が分けられないSharp in Time, Blurred in Frequency
只听很短的一瞬,能说准那一声咔嗒落在什么时候。但这么短的一段里音高本来就听不出来,泛音挤成一团,共鸣落在哪一带也说不清。ごく短い一瞬だけを聞けば、そのカチッという音がいつ鳴ったかは正確に言えます。しかしこれほど短い区間では音の高さがそもそも聞き取れず、倍音は一塊になり、どの帯域で共鳴しているかも分かりません。Listen to only a brief instant and you can say exactly when the click occurred. But in so short a stretch the pitch cannot be made out at all: overtones crowd together and there is no telling which band is resonating.
频率分得清,时间就分不准周波数を細かく分けると、時間が分けられないSharp in Frequency, Blurred in Time
连着听很长一段,音高与共鸣能分得很细,但短促的撞击被摊平:响在哪一刻、响了几次、衰减多快全部丢失,而这最能说明装配是不是松了。長く続けて聞けば、音の高さと共鳴は細かく聞き分けられます。しかし短い衝突音はならされてしまい、いつ・何回鳴ったか、どれだけ速く減衰したかが失われます。組付けの緩みを最もよく表すのは、まさにその部分です。Listen across a long stretch and pitch and resonance separate finely, but a brief impact is flattened out: when it struck, how many times, how fast it decayed — all lost, and those are exactly what reveal a loose assembly.
本引擎的做法:几种尺度同时用,信谁由模型学出本エンジンの方法:複数のスケールを同時に使い、どれを信じるかはモデルが学ぶOur approach: several scales at once, and the model learns which to trust
粗细不同的尺度,同时用粗さの異なるスケールを同時に使うSeveral Scales at Once
不二选一,而是对同一段声音同时用几种跨度不同的尺度各分辨一遍。长跨度把音高与共鸣分得稳,短跨度把敲击时刻与衰减分得准,两类信息都在手。どちらかを選ぶのではなく、同じ音を長さの異なる複数のスケールで同時に捉えます。長いスケールは音の高さと共鳴を安定して捉え、短いスケールは打撃の瞬間と減衰を正確に捉え、両方の情報が手元に残ります。Rather than choosing, the same sound is resolved at once on several spans. The longer spans hold pitch and resonance steady; the shorter ones pin down the moment of impact and its decay, and both stay available.
让模型在训练中学会该信哪一种尺度どのスケールを信じるかを学習の中でモデル自身が身につけるThe Model Learns in Training Which Scale to Trust
敲击的瞬间该信短跨度,持续的啸叫该信长跨度。这不是人工调的规则,而是模型在训练中从正常声音里自己学出来的,不必逐工位手调。打撃の瞬間は短いスケールを、持続する異音は長いスケールを信じるべきです。これは人が調整した規則ではなく、モデルが学習の中で正常音から自ら身につけたものであり、工程ごとに手で調整する必要はありません。At the instant of an impact the shorter spans should be trusted; for a sustained squeal, the longer ones. This is not a hand-tuned rule: the model learns it for itself from normal sound during training, so nothing has to be adjusted station by station.
同一处声音在不同尺度下怎么变同じ箇所の音がスケールによってどう変わるかに着目するHow the Same Sound Changes Across Spans
模型还关注同一处声音,从长跨度换到短跨度时怎么变。撞击类跨度越短越突出,持续类在各种尺度下都稳定,这种变化本身就是最直接的线索,也被学进表征里。モデルは同じ箇所の音が、長いスケールから短いスケールへどう変わるかにも着目します。衝突音はスケールが短いほど際立ち、持続音はどのスケールでも安定します。この変わり方そのものが最も直接的な手がかりであり、表現の中に学び取られます。The model also attends to how the same sound changes as the span shortens. An impact stands out more as the span shortens; a sustained sound stays steady across every span. That pattern of change is itself the most direct clue, and it too is learned into the representation.
学得会:自监督预训练学び取る:自己教師あり学習Self-Supervised Pre-Training
基础能力来自自监督预训练,不来自人工标注的缺陷样本。基礎能力は自己教師あり学習による事前学習から得られ、欠陥のラベルには依存しません。Core ability comes from self-supervised pre-training, not from labelled defects.
客户最常问的是要提供多少个不良品。本引擎的核心是一个深度神经网络,先在大量无标注的工业声音上做自监督预训练,学正常的声音长什么样,不依赖缺陷标注。
最初のご質問は、不良品をどれだけ用意すればよいかというものです。本エンジンの中核はディープニューラルネットワークであり、まずラベルなしの産業音を大量に用いた自己教師あり学習で事前学習を行い、正常な音の姿を学びます。欠陥のラベルには依存しません。
The first question customers ask is how many defective parts to supply. At the core of the engine is a deep neural network, pre-trained self-supervised on a large body of unlabelled industrial sound to learn what normal sounds like, without relying on defect labels.
自监督预训练的循环自己教師あり学習による事前学習の循環The Self-Supervised Pre-Training Loop
-
01
听聴くListen
送进一段现场声音,网络形成自己的表征。
現場の音を入力すると、ネットワークは自らの表現を形づくります。
A segment of line sound goes in, and the network forms its own representation of it.
-
02
还原復元するReconstruct
再要求模型只凭这份表征重新描绘原声,不许再回头听原声。
次に、その表現だけを頼りに元の音を描き出させます。元の音は参照できません。
It must then redraw the original from that representation alone, without hearing it again.
-
03
自查自己点検するCompare
把还原结果与原声逐处对照,差在哪里模型自己就能分辨。
復元結果と元の音を照合します。ずれはモデル自身が把握できます。
The reconstruction is compared with the original; the model can tell where the gaps are.
-
04
修正修正するCorrect
对不上的地方成为改进方向,回到第一步再来一遍。
合わなかった箇所を改善の方向として、最初の工程に戻ります。
The mismatches become the direction for improvement, and the cycle starts again.
为什么这样能学到有用的东西なぜこれで有用なものが身につくのかWhy This Teaches Something Useful
学出来的强过人设计的学び取った表現は人が設計した指標に勝りますLearned Beats Hand-Designed
传统方法靠人决定该分辨哪些指标,人想不到的规律就永远进不了判据。要准确还原设备声音,网络必须自己学出谐波成组、包络起伏、共振抬升与噪声底质地,也包括人没想到的线索。
従来手法では、どの指標を手がかりにするかを人が決めます。人が思いつかなかった規則性は判定根拠に入りません。設備の音を正確に復元するには、倍音のまとまり、包絡の起伏、共振の持ち上がり、暗騒音の質感を、ネットワーク自身が学び取る必要があります。人が思いつかなかった手がかりも含まれます。
In the traditional approach a person decides which indicators to rely on, so any regularity nobody thought of never enters the criteria. To redraw a machine sound faithfully, the network must learn harmonic families, envelope shape, resonant lift and noise-floor texture for itself — including cues no one thought to specify.
学会常态,才谈得上察觉异常正常を知って初めて異常に気づけますKnow Normal Before You Notice Abnormal
异响之所以被察觉,是因为它偏离了本该如此的样子。常态学扎实,偏离就有了参照,没见过的缺陷也能察觉。
異音に気づけるのは、本来あるべき姿から外れているからです。正常をしっかり学べばずれの基準が得られ、見たことのない欠陥でも捉えられます。
An abnormal noise stands out because it departs from what it should have been. Once normal is learned, that departure has a reference, even for an unseen defect.
一个类比たとえ話An Analogy
能把刚听过的旋律准确复述出来,说明真的听懂了它的结构,而不是背下了一个曲名标签。只会贴标签的人,换一首没听过的曲子就无从下手。
聞いたばかりの旋律を正確に口ずさめるのは、その構造を本当に理解している証です。曲名というラベルを覚えただけの人は、初めての曲になると手がかりを失います。
Humming back a melody you just heard shows you grasped its structure, not merely its title. Someone who only memorises titles is helpless with an unfamiliar tune.
这对现场意味着什么現場にとっての意味What This Means on the Line
起步不依赖大量缺陷样本導入時に大量の不良サンプルは不要ですGetting Started Does Not Hinge on Defect Samples
现场的不良品稀少、种类零散,也拿不到成套标注。导入阶段主要需要能代表当前工况的正常声音。
現場の不良品は少なく種類もばらばらで、体系的なラベルは揃いません。導入段階で必要なのは現在の工程条件を代表する正常音です。
Defective parts are rare, scattered across types, and rarely fully labelled. The introduction phase mainly needs normal sound representing current conditions.
预训练打底,现场适配收口事前学習で土台を、現場適応で仕上げをPre-Training First, On-Site Adaptation After
预训练在广泛的工业设备声音上学到通用表征,再通过现场适配对准每条产线,因此不易被一条线的偶然特征带偏。
事前学習では幅広い産業設備の音から汎用的な表現を学び、その後の現場適応で各ラインに合わせ込みます。そのため特定ラインの偶発的な特徴に引きずられにくくなります。
Pre-training learns a general representation across a wide range of industrial machine sound; on-site adaptation then fits it to each line, so it is not pulled off course by one line's quirks.
无标注不等于不需要现场信息ラベルなしでも現場情報は必要ですLabel-Free Does Not Mean Context-Free
预训练给的是基础听觉能力;要在具体产线上判定,仍需现场常态声音作参照和稳定的采集条件。
事前学習が担うのは基礎的な聴取能力です。個々のラインで判定を出すには、現場の正常音という参照と、安定した収録条件が引き続き必要です。
Pre-training provides the basic listening ability. Judging on a specific line still needs that line's normal sound as a reference and stable recording conditions.
辨得全:三个互补视角取りこぼさない:三つの相補的な視点Covering the Whole: Three Complementary Views
同一次推理的表征用三种方式读出,同一个缺陷在不同视角下的可见性差别很大。一度の推論で得た表現を三通りに読み出します。同じ欠陥でも視点により捉えやすさが大きく異なります。One inference, read three ways — a defect obvious in one view can be invisible in another.
三个视角读的是同一个深度神经网络的内部表征,不是三套人工设计的特征。同一份表征,读出方式不同,能分辨出的东西就不同;这也是它们能共享一次推理、且互补而不重复的根本原因。
三つの視点が読み出しているのは、同一のディープニューラルネットワークの内部表現であり、人が設計した三種類の特徴量ではありません。同じ表現でも読み出し方が異なれば捉えられるものは変わります。三つが一度の推論を共有でき、しかも重複ではなく相補となるのは、このためです。
All three views are read from the internal representation of one and the same deep neural network, not from three sets of hand-designed features. The same representation, read differently, tells apart different things — which is why they can share a single inference and still complement rather than repeat each other.
三个视角各自分辨什么三つの視点はそれぞれ何を捉えているかWhat Each View Picks Up
整体声纹视角全体音紋の視点Overall Sound Signature
像给这段声音拍一张全景照。它最稳定,冷启动阶段最先可以信任。
その音の全景写真を一枚撮るようなものです。最も安定しており、立ち上げ初期に最初に信頼できる経路です。
A wide shot of the segment. It is the steadiest path, and the first one to trust during a cold start.
听感对齐视角聴感に沿った視点Perceptually Aligned View
把同一份表征按与人耳相关的方向重排后再比较,相当于换一把刻度更合适的尺子。
同じ表現を人の耳に関わる方向に並べ直してから比較します。目盛りのより適した物差しに持ち替えるようなものです。
The same representation is rearranged along directions tied to hearing, then compared — like swapping in a ruler with a better-suited scale.
时间纹理视角時間的な質感の視点Temporal Texture
整体概括会把起伏抹平,这一路专门把节奏与周期性保留下来。
全体をまとめる読み出しでは起伏がならされるため、この経路はリズムと周期性を意図的に残します。
A summary smooths movement away; this path deliberately keeps the rhythm and any periodicity.
| 视角視点View | 擅长回答得意な問いAnswers best | 直觉解释直感的な説明Intuition |
|---|---|---|
| 整体声纹全体音紋Overall Signature | 整体是否偏离常态全体として正常から外れていますかDrifted from normal as a whole? | 全身照:体型变化一眼可见,小痣看不清全身写真です。体格の変化は一目で分かり、細かな点は見えませんA full-length photo: build is obvious, a small mark is not |
| 听感对齐聴感に沿った視点Perceptually Aligned | 人听起来刺不刺耳人の耳に耳障りに聞こえますかWould a person hear it as harsh? | 同一个东西,换一把刻度更合适的尺子同じ対象を、より適した目盛りの物差しで測りますThe same object, measured with a better-suited scale |
| 时间纹理時間的な質感Temporal Texture | 持续还是一下一下継続的ですか、断続的ですかContinuous, or in bursts? | 不是照片,是录像的节奏写真ではなく、動画のリズムですNot a photograph, but the rhythm of a video |
互补,不是冗余冗長ではなく、相補ですComplementary, Not Redundant
有的缺陷在整体概括下几乎分辨不出,换到时间纹理却一目了然,反过来也有。哪个视角显形事先无法预判,所以三路同时保留。
全体をまとめた読み出しではほとんど差が出ないのに、時間的な質感では明らかになる欠陥があり、その逆もあります。どの視点で表に出るかは事前に予測できないため、三つの経路を同時に保持しています。
Some defects barely register in the overall summary yet stand out plainly in temporal texture, and the reverse happens too. Since no one can predict which view will reveal a given defect, all three stay open.
代价与前提コストと前提Cost and Preconditions
三路共享一次推理,但各自算分三経路は推論を共有し、評価は個別に行いますOne Shared Inference, Three Separate Scores
三条路径读的是同一次推理,增加的只是读出环节,对节拍影响很小。三者量纲与稳定性不同,混成一串会让量纲大的一路主导,因此各自独立评分与校准,再在分数层面汇总。
三つの経路は同一の推論を読み出しており、増えるのは読み出しの工程だけで、タクトへの影響は小さく済みます。三者は尺度も安定性も異なり、一本につなぐと尺度の大きい経路が結果を支配するため、それぞれ独立に評価・校正したうえで、スコアの段階でまとめます。
All three read the same inference, so only the readout step is added, with little effect on cycle time. They differ in scale and stability, and strung together the largest would dominate, so each is scored and calibrated on its own and combined at the score level.
对得上人耳:听感对齐人の耳に合わせる:聴感整合Aligned With the Human Ear
判定异响靠的不是能量大小,而是粗糙感、波动感、尖锐感。異音の判定は音の大きさではなく、粗さ・変動感・鋭さに依存します。Judging abnormal noise depends on roughness, fluctuation and sharpness, not on level.
现场判定异响的最终裁判是人,判据不是能量表上的读数。引擎必须在人听起来怎么样这一层上对齐。
現場で異音を最終的に判定するのは人であり、その基準はレベル計の読み値ではありません。エンジンは、人にどう聞こえるかという層で整合をとる必要があります。
On the line the final judge is a person, not a reading on a level meter. The engine has to align at the layer of how something sounds.
两段一样响的声音同じ大きさの二つの音Two Sounds, Equally Loud
两段声音能量完全相同,人听起来却可以一个平稳、一个刺耳。区别在于能量如何分布、如何起伏、有无明显的调制。
エネルギーが等しい二つの音でも、一方は落ち着いて、もう一方は耳障りに聞こえます。違いは、エネルギーがどう分布し、どう変動し、変調があるかどうかにあります。
Two sounds can carry identical energy, one calm and one harsh. The difference lies in how that energy is spread, how it fluctuates, and whether it is modulated.
人耳在意的是哪几件事人の耳が気にしているものWhat the Ear Actually Cares About
这些是声学领域公开的听感属性,引擎用作参照方向。
これらは音響分野で公知の聴感属性であり、エンジンは参照の方向として利用します。
These are publicly known perceptual attributes, used by the engine as reference directions.
引擎怎么用这些属性エンジンはこれらの属性をどう使うかHow the Engine Uses These Attributes
当参照方向,让放大有的放矢参照の方向として用い、的を絞って拡大しますA Reference Direction, So Magnification Is Targeted
引擎不是先算出这些属性再判好坏,而是用它们指引方向,在深度神经网络学出的表征里挑出与听感相关的成分,放到同一量级上比较。放大所有细微成分会让噪声一起变大,所以要挑选。
エンジンは属性をそのまま良否判定に用いず、方向の手がかりとして使い、ディープニューラルネットワークが学び取った表現から聴感に関わる成分を選び出して同じ尺度に揃えます。すべてを拡大すればノイズも大きくなるため、選び出すことが必要です。
The engine does not judge from these attributes directly; it uses them as a guide to pick out the hearing-related components inside the representation the deep neural network has learned and bring them onto a common scale. Magnify everything indiscriminately and the noise grows too, so the work is selection.
不是创造信息,是让信息变得可读情報を生み出すのではなく、既にある情報を読めるようにしますIt Makes Existing Information Legible
现场声音的差异绝大部分来自背景与工况:温度、批次、周边工位。缺陷带来的那部分小得多,直接比较会被完全盖住,就像在大秤上称一根羽毛——秤没坏,只是刻度不合适。
現場の音の違いの大部分は背景と工程条件に由来します。温度、ロット差、周辺工程などです。欠陥による違いははるかに小さく、そのまま比較すると覆い隠されます。大きな秤で羽根を量るようなもので、秤ではなく目盛りが合っていないだけです。
Most of the difference between sounds on a line comes from background and conditions: temperature, lot variation, neighbouring stations. The defect's share is far smaller and is buried in a direct comparison, like weighing a feather on a heavy-duty scale. The scale is not broken; its graduations do not suit the task.
对齐听感相当于换一把刻度更合适的尺子:数据没变,被掩盖的细微差异却回到了可以分辨的尺度上。
聴感に合わせることは、より適した目盛りの物差しに持ち替えることに相当します。データは変わらないまま、覆い隠されていた違いが識別できる尺度に戻ります。
Aligning with perception amounts to picking up a better-suited ruler: the data is unchanged, yet masked differences return to a scale where they can be told apart.
这把尺子要按现场来定この物差しは現場に合わせて決めますThe Ruler Must Be Fitted On Site
刻度依据当前产线、当前工况的正常声音,通过现场适配确定。换线、换工位或工况明显变化,都要重新对刻度。
この層の目盛りは、現在のラインと工程条件の正常音に基づき、現場適応によって決めます。ライン変更や条件の大きな変化があれば、取り直しが必要です。
The graduations here are set by on-site adaptation from normal sound on the current line under current conditions. A change of line, station or conditions calls for re-setting them.
判得准:分数层融合与按线择优見極める:スコア層での統合とライン別の最適化Judging Reliably: Score-Level Fusion and Per-Line Selection
三个视角各自成为完整判断,在分数层汇合,最终结论落在工件上。三つの視点がそれぞれ独立した判断を出し、スコア層で統合され、結論はワーク単位で下されます。Each viewpoint forms a complete judgment, the three meet at the score level, and the verdict lands on the part.
三个视角有了之后,问题变成怎么把三份意见合成一个结论。就像三位老师傅分别听同一件产品,不该把耳朵缝在一起,而应各自下结论再汇总。 三つの視点が揃うと、次は三つの意見をどう一つの結論にまとめるかが課題になります。三人の熟練者が同じ製品を聞くとき、耳を縫い合わせるのではなく、各自が結論を出して持ち寄るのと同じです。 Once three viewpoints exist, the question becomes how to turn three opinions into one conclusion. Three experienced inspectors listening to the same part should not have their ears stitched together; each reaches a conclusion, and the opinions are pooled.
不要硬拼単純な連結はしないDo Not Concatenate
三个视角的数值尺度本来就不同,硬拼在一起会被跨度最大的那个视角主导,另两个视角的细微差异被压成噪声。这就像把体温、血压、听力的读数直接相加,得先各自换算成偏离健康人群多少,再谈综合。 三つの視点は数値の尺度が元々異なり、単純に連結すると幅の最も大きい視点に支配され、他の二つの微細な差はノイズに埋もれます。体温・血圧・聴力の測定値をそのまま合計するようなもので、まず各項目を健常な集団からの隔たりに換算してから総合を語るべきです。 The three viewpoints differ in scale, so concatenating them lets the widest-ranging one dominate while the subtle differences of the other two flatten into noise. It is like summing body temperature, blood pressure, and a hearing test: each must first be converted into how far it deviates from a healthy population.
把表征拼成一条表現を一本に連結するConcatenate the Representations
强势视角说了算,弱势视角被淹没,出问题也难以定位到是哪一个视角。値の大きい視点が判断を支配し、他の視点は埋もれます。不調の原因がどの視点にあるのかも切り分けにくくなります。The dominant viewpoint decides everything, the quieter ones are drowned out, and trouble is hard to trace back to a single viewpoint.
不采用採用しないNot Used
各自打分,再合意见個別に採点し、意見を統合するScore Separately, Then Combine
每个视角独立回答离本线正常有多远,换算成可比的相对位置后在分数层汇总,各自的分辨力都保留下来。各視点が当該ラインの正常からどれだけ離れているかを独立して答え、比較可能な相対位置に換算したうえでスコア層で統合します。各視点の分解能はそのまま保たれます。Each viewpoint answers on its own how far the part sits from normal on this line, and the answers are converted into comparable positions before being combined at the score level, so every viewpoint keeps its resolving power.
本引擎采用本エンジンの方式Our Approach
按产线择优ライン別の最適化Per-Line Selection
不同产品、不同工位,主导的缺陷类型本来就不同:有的表现为整体音色偏移,有的表现为时间上的细碎起伏。所以给出统一机制,让每条线用自己的证据挑选该倚重的视角。分数来自模型推理,不是人工阈值;人工设定的只有误判预算一个旋钮。 製品や工程が変われば、主となる不良のあらわれ方も変わります。音色全体のずれとして現れる場合もあれば、時間方向の細かな揺らぎとして現れる場合もあります。そこで統一された仕組みを用意し、各ラインが自らの証拠で重視すべき視点を選べるようにしています。スコアはモデルの推論から得られるもので、人手で決めた閾値ではありません。人が設定するのは、誤判定の許容量という一つのつまみだけです。 The dominant defect type differs by product and by station: on one line it is a shift in overall timbre, on another it is fine fluctuation over time. So the engine provides one common mechanism and lets each line use its own evidence to choose the viewpoints it relies on. The scores come from model inference, not from hand-set thresholds; the only knob left to people is the false-call budget.
一件产品在采音期间被切成若干段分别打分,任一段异常即判该件异常。异常常常只出现在一瞬间,不应被整段声音平均掉。一つの製品は集音中にいくつかの区間に分けて採点され、いずれか一つでも異常と判定されればワークは異常となります。異常は一瞬だけ現れることが多く、音全体で平均されて消えてはなりません。During recording a part is divided into several clips scored separately, and if any one clip is abnormal the part is abnormal. A fault often appears for only an instant and must not be averaged away.
跟得上:滚动校准与误判预算追従する:ローリング校正と過検出バジェットKeeping Pace: Rolling Calibration and the False-Alarm Budget
门槛随近期正常件持续重定;品质与产能的取舍由客户用一个旋钮设定。しきい値は直近の正常品に合わせて定め直され、品質と生産性のトレードオフはお客様が一つのつまみで設定します。The threshold is re-derived from recent good parts, and the quality-throughput trade-off is set by the customer with one dial.
产线是活的:批次、季节、工装磨损都会让声音的基线慢慢移动。出厂时定死的固定门槛,久了必然要么太松要么太紧。 ラインは生きています。ロット、季節、治具の摩耗によって音の基準はゆっくり移動します。出荷時に固定したしきい値は、いずれ緩すぎるか厳しすぎるかになります。 A production line is a living thing: lots, seasons, and fixture wear all shift the baseline of the sound. A threshold fixed at the factory will sooner or later be too loose or too tight.
门槛来自近期的正常件しきい値は直近の正常品から決まるThe Threshold Comes from Recent Good Parts
引擎持续收集本线近期正常件的分数,汇成现在的正常分布,门槛落在其上并按天重新推定。批次与季节的缓慢变化因此被自然吸收。不是拿出厂的尺子量所有年份的产品,而是每天早上先量今天的正常件,再决定今天的界线。校准只重定门槛,预训练的深度听觉模型与其模型版本保持不变。 エンジンは当該ラインの直近の正常品のスコアを集めて今の正常分布を作り、しきい値はその上に置かれ日次で推定し直されます。ロットや季節による緩やかな変化は自然に吸収されます。出荷時のものさしで何年分もの製品を測るのではなく、毎朝まず今日の正常品を測り、今日の線引きを決めるのです。校正で定め直すのはしきい値だけで、事前学習済みの深層聴覚モデルとそのモデルバージョンは変わりません。 The engine gathers the scores of the line's recent good parts into a distribution of normal, and the threshold sits on it and is re-derived each day. Slow drift from lots and seasons is absorbed naturally. Rather than measuring every year's production with a ruler fixed at the factory, the line measures today's good parts each morning and then decides today's boundary. Calibration re-derives only the threshold; the pre-trained deep listening model and its model version stay unchanged.
误判预算:一个旋钮過検出バジェット:一つのつまみThe False-Alarm Budget: One Dial
漏检和误判无法同时降到最低,这是取舍,不是缺陷。由客户设定能接受多少比例的正常件被拦下复核,系统据此反推当日门槛。拧严一点复核量上升,拧松一点产能轻松、边缘异常更可能通过。 見逃しと過検出を同時に最小にすることはできません。これは欠陥ではなくトレードオフです。正常品のうちどれだけの割合が再検査に回ることを許容するかをお客様が設定し、そこから当日のしきい値を逆算します。厳しくすれば再検査の工数が増え、緩くすれば生産は楽になりますが境界付近の異常が通過しやすくなります。 Missed defects and false alarms cannot both be minimized; that is a trade-off, not a flaw. You set what share of good parts you accept being held for re-inspection, and the day's threshold is derived from it. Tighten it and workload rises; loosen it and throughput eases while borderline faults are likelier to pass.
可调調整できるAdjustable
不同工位、不同阶段可以设不同的严格程度。试产期与量产期本就不该一样。工程ごと、段階ごとに異なる厳しさを設定できます。試作期と量産期が同じである必要はありません。Different stations and phases can be set to different strictness. Pilot and mass production need not share one setting.
可解释説明できるExplainable
门槛对应的是复核工作量,是一件件要人确认的产品,品质会议上可以直接讨论。しきい値は再検査の工数、すなわち人が一つずつ確認する製品の数に対応し、品質会議でそのまま議論できます。The threshold corresponds to re-inspection workload, to parts a person confirms one by one, and can be discussed directly in a quality review.
可复盘振り返れるReviewable
设定值与实际拦下比例的偏离是健康度信号。偏离持续拉大,先查采音条件与工艺变更,而不是先改门槛。設定値と実際に止められた割合とのずれは健全性の信号です。ずれが拡大し続ける場合は、しきい値を変える前に集音条件と工程変更をご確認ください。The gap between the setting and the share actually held is a health signal. If it keeps widening, check recording conditions and process changes before touching the threshold.
复核回流:闭环再検査結果のフィードバック:クローズドループReview Results Flow Back: Closing the Loop
被拦下的工件由现场人工复核,结论回流成为确认样本。系统因此逐步知道本线真正的异常听起来是什么样。 止められたワークは現場で人が再検査し、その結論が確認済みサンプルとして戻されます。システムはこのラインで実際の異常がどう聞こえるかを少しずつ把握していきます。 Held parts are re-inspected by people on the floor, and their conclusions flow back as confirmed samples. The system thus comes to learn what a real fault on this line sounds like.
回流的主要是被系统拦下的工件,学到的异常样貌因此受复核范围与质量影响。一开始就没被拦下的异常,很难出现在学习材料里。建议导入初期对放行件也抽检,结论一并回流。フィードバックされるのは主にシステムが止めたワークですので、学ばれる異常の姿は再検査の範囲と質に左右されます。最初から止められなかった異常は、学習材料に現れにくくなります。導入初期は合格として流した品にも抜き取り検査を行い、その結果も併せてフィードバックしてください。What flows back is mainly the parts the system held, so the picture of abnormality it learns depends on the scope and quality of re-inspection. A fault never held rarely appears in the learning material at all. In the early phase, sample passed parts as well and feed those conclusions back too.
用得稳:现场前提与适用边界安定して使う:現場の前提と適用範囲Running Stably: Site Prerequisites and Scope of Application
系统能否稳定工作,一半在算法,一半在采音与工况守不守得住。安定して機能するかどうかは、半分がアルゴリズム、半分が集音と工況を保てるかどうかで決まります。Stable operation depends half on the algorithm and half on whether recording and operating conditions hold steady.
任何声学检测都建立在一个前提上:这一次听到的声音和上一次可比。把边界说清楚,比宣称无所不能更有价值。 あらゆる音響検査は、今回聞いた音と前回聞いた音が比較可能であるという前提の上に成り立ちます。適用範囲を明示することは、万能であると主張するより価値があると考えます。 Every acoustic inspection rests on one premise: what is heard this time is comparable with what was heard last time. Stating the limits is more useful than claiming there are none.
| 前提前提Prerequisite | 为什么重要なぜ重要かWhy It Matters | 现场做法現場での対応What to Do on Site |
|---|---|---|
| 采音一致性集音の一貫性Consistent Recording Setup | 麦克风、位置、指向、增益、工装任一变更,绝对水平与音色都会移位。マイク、位置、指向、ゲイン、治具のいずれかが変われば、音の絶対レベルと音色が移動します。A change of microphone, placement, orientation, gain, or fixture shifts both absolute level and timbre. | 视同新部署,重新校准并重建正常基线,变更内容与日期留档。新規導入と同等に扱い、再校正と正常基準の再構築を行い、変更内容と日付を記録してください。Treat it as a new deployment: recalibrate, rebuild the baseline, and record what changed and when. |
| 环境与工况一致性環境と工況の一貫性Stable Noise and Operating Conditions | 正常范围内的背景波动可以吸收,节拍变化与新增噪声源不行。通常の範囲の暗騒音の変動は吸収できますが、タクトの変更や新たな騒音源は吸収できません。Ordinary background fluctuation can be absorbed; a change of cycle time or a new noise source cannot. | 保持节拍与背景条件稳定,出现明显变化即重新验证。タクトと背景条件を安定に保ち、明らかな変化があれば再検証してください。Keep cycle and background steady; revalidate whenever a clear change occurs. |
| 片段与节拍匹配集音区間と工程タクトの整合Clips Aligned to the Process Cycle | 切得太短会把特征切碎,切得太长会把瞬时异常平均掉。短すぎると特徴が分断され、長すぎると瞬間的な異常が平均化されて消えます。Clips too short fragment the evidence; clips too long average away momentary faults. | 导入阶段与工艺方共同确定触发方式与片段口径。導入段階で工程担当と共に集音のトリガ方法と区間の切り方を決めてください。Agree the recording trigger and the clip boundaries with process engineering during deployment. |
| 基线积累与现场适配基準の蓄積と現場適応Baseline Accumulation and On-Site Adaptation | 现场适配用正常件把预训练模型对齐到本产线,基线不足时判定易被偶然波动带偏。現場適応では正常品を用いて事前学習済みモデルを本ラインに合わせます。基準が不十分なうちは、判定が偶発的なばらつきに引きずられます。On-site adaptation uses good parts to align the pre-trained model to this line; while the baseline is thin, chance variation pulls decisions off course. | 导入初期先积累,覆盖不同班次后完成现场适配再启用;模型版本固定,同一版本的推理结果可复现,新产线靠现场适配接入而不是重写规则。導入初期はまず蓄積し、異なる直を網羅したうえで現場適応を完了してから稼働してください。モデルバージョンは固定され、同一バージョンでの推論結果は再現できます。新しいラインはルールを書き直すのではなく、現場適応によって立ち上げます。Accumulate first, cover the different shifts, complete on-site adaptation, then go live. The model version is fixed and inference is reproducible under the same version; a new line is brought up by on-site adaptation rather than by rewriting rules. |
| 判定口径与复核判定基準と再検査Shared Criteria and Review | 上下游对什么算异常的定义不一致,回流样本就会互相矛盾。何を異常とするかの定義が前後工程で揃わなければ、戻されるサンプルが互いに矛盾します。If upstream and downstream disagree on what counts as abnormal, the samples fed back contradict one another. | 明确复核标准与责任岗位,各班组用同一份标准。再検査の判定基準と責任者を明確にし、どの班も同一の基準をご使用ください。Define the criteria and the responsible role, and have every shift use the same standard. |
导入方式:影子期導入方法:シャドー運用期間Deployment: The Shadow Period
并行运行,只记录不干预並行稼働し、記録のみで介入しないRun in Parallel, Record Only
引擎与现有流程同时工作,输出只作记录,不影响放行。エンジンは既存の工程と並行して稼働し、出力は記録のみで合否には影響しません。The engine runs alongside the existing process; its output is recorded only and does not affect release.
积累与比对蓄積と突き合わせAccumulate and Compare
形成正常基线并完成现场适配,同时与现场判定逐件比对,确认口径一致。正常基準を形成して現場適応を完了しつつ、現場の判定と一件ずつ突き合わせ、基準の一致を確認します。The baseline takes shape and on-site adaptation completes, while decisions are compared part by part with the site's own to confirm shared criteria.
切入判定判定への切り替えSwitch Into Live Judgment
客户认可后按设定的误判预算参与判定,随时可切回只记录。お客様のご承認後、設定された過検出バジェットに従って判定に参加し、いつでも記録のみへ戻せます。After the customer approves, it judges under the agreed false-alarm budget, and can return to record-only at any time.
影子期真正的价值,是把什么算异常在系统与现场之间对齐。 シャドー運用期間の本当の価値は、何を異常とするかをシステムと現場で揃えることにあります。 The real value of the shadow period is aligning the system and the site on what counts as abnormal.
适用边界適用の範囲Scope and Limits
以上任何一项前提发生变更,都应先重新校准并做一段影子期复核,再恢复正式判定,让判定始终建立在可比的声音之上。上記の前提のいずれかが変更された場合は、再校正と一定期間のシャドー運用による確認を行ってから正式な判定に戻し、判定を常に比較可能な音の上に成り立たせてください。If any prerequisite above changes, recalibrate and run a shadow review before returning to live judgment, so every decision keeps resting on comparable audio.
常见问题よくあるご質問Frequently Asked Questions
导入前最常被问到的问题,按实况作答,不作效果承诺。導入前に最も多くいただくご質問に実情に即してお答えします。効果のお約束はいたしません。The questions we are asked most often before deployment, answered as things stand, with no promises of performance.
以下问题来自现场的实际关切;真实表现请在贵司产线上以影子期验证。 以下のご質問は現場から実際に寄せられたものです。実際の性能は、貴社ご自身のラインでシャドー運用期間を通じてご確認ください。 These questions come from real concerns on the floor. Actual behavior should be verified on your own line, through a shadow period.
传统方法由人先选定指标、逐个设门槛,人想不到的异常进不了判据。本引擎是深度学习系统,在大量无标注声音上预训练,自己学出该分辨什么,即表征(模型自己形成的内部表示),再按产线现场适配,因此能听出能量没变、音色变了这类差异。従来手法では、どの指標を手がかりにするかを人があらかじめ決め、指標ごとにしきい値を設けますので、人が想定できない異常は判定基準に入りません。本エンジンは深層学習のシステムで、大量のラベルなし音で事前学習し、何を手がかりにすべきかを自ら学びます。これが表現、すなわちモデルが自ら形成する内部的な表し方であり、その上でラインごとに現場適応します。そのため、エネルギーは変わらず音色だけが変わったような違いも捉えられます。Conventional methods have people decide in advance which indicators to rely on and set a threshold on each, so an anomaly nobody anticipated never enters the criteria. This engine is a deep learning system: pre-trained on large amounts of unlabelled sound, it learns for itself what to listen for, its representation, and is then adapted on site to the line. That is how it tells apart differences such as timbre changing while energy does not.
主体学习是自监督的,先学会正常声音长什么样,不依赖缺陷标注即可开始。少量已确认的异常件有助于对齐口径,但不是前提。学習の主体は自己教師ありで、まず正常な音がどのようなものかを学びますので、不良ラベルがなくても開始できます。確認済みの異常品が少量あれば基準合わせに役立ちますが、前提ではありません。The main learning is self-supervised: it first learns what normal sound looks like, so it can start without labeled defects. A few confirmed abnormal parts help align criteria, but are not a precondition.
门槛按近期正常件持续重定,缓慢变化自然吸收;工艺、工装或采音条件的实质变更应重新校准并做影子期复核。しきい値は直近の正常品から継続的に定め直され、緩やかな変化は自然に吸収されます。工程、治具、集音条件の実質的な変更があった場合は、再校正とシャドー運用による確認を行ってください。The threshold is re-derived continuously from recent good parts, so slow change is absorbed naturally. A substantive change to process, fixture, or recording setup calls for recalibration and a shadow review.
每次推理都保留采音、片段与分数,可逐件回溯,判错的件回流校准;最终判定权仍在现场。推論ごとに集音、区間、スコアを保持し、一件ずつ遡って確認でき、誤判定は校正に反映されます。最終的な判定権限は現場にあります。Every inference keeps its recording, clip, and score for review, and a wrong call feeds back into calibration; the final say remains with the site.
推理在采音完成后进行,目标是在工位节拍内完成;边缘端还是服务端按项目条件确定。推論は集音の完了後に行われ、工程のタクト内で完了することを目標とします。エッジ側かサーバ側かはプロジェクトの条件で決めます。Inference runs after recording finishes, with the goal of completing within the station's cycle; edge or server is decided by project conditions.
可以回溯到哪一段触发、哪个视角偏离最显著,并调出那段声音收听;这是指向证据而非因果解释。どの区間が引き金となり、どの視点の逸脱が最も大きかったかまで遡ることができ、該当する音をその場で再生できます。これは証拠を指し示すものであり原因の説明ではなく、原因の判断は現場に委ねられます。You can trace which clip triggered the flag and which viewpoint deviated most, and play back that audio. This points to evidence rather than cause; the cause is still judged on the floor.
流程固定:接入采音、积累基线、影子期验证、切入判定,时间取决于产量与异常频度。手順は共通で、集音の接続、正常品の基準の蓄積、シャドー運用での検証、判定への切り替えとなります。所要期間は生産量と異常の発生頻度によって決まります。The sequence is fixed: connect recording, accumulate the baseline, validate in shadow, then switch into live judgment. How long it takes depends on output and how often faults occur.
以工件标识关联,输出判定结论与分数供上位系统记录统计,接口与触发按项目约定。ワーク識別子で紐づけ、判定結果とスコアを出力し、上位システムでの記録・集計にご利用いただけます。インタフェースの形式とトリガ方法はプロジェクトで取り決めます。Results are linked by part identifier, and the decision and score are output for the host system to record and aggregate. Interface format and trigger are agreed within the project.
底层的深度神经网络保持稳定,升级只新增读出方式。新的模型版本可并行影子验证,确认一致后再切换,必要时回退。基盤となるディープニューラルネットワークは安定して維持され、更新は読み出し方の追加にとどまります。新しいモデルバージョンは並行してシャドー検証し、一致を確認してから切り替え、必要に応じて切り戻せます。The underlying deep neural network stays stable; an update only adds new ways of reading it out. A new model version can be shadow-validated in parallel, switched over after agreement, and rolled back if needed.
具体产线的可行性,建议以一次影子期作为共同评估依据;贵司产品上的结论比通用说明更有参考。具体的なラインでの適用可否は、一度のシャドー運用を共通の評価根拠とすることをお勧めします。貴社の製品から得られた結論のほうが、一般的な説明より参考になります。For feasibility on a specific line, we suggest a shadow period as the shared basis; a conclusion from your own products beats any general description.