The call came from Nina.
By then she chaired the regional driver advisory group and enjoyed reminding me that I had escaped before the meetings became worse.
“You retired at the right time.”
“I did not retire.”
“From committees.”
“That counts.”
“Need your opinion.”
“No.”
“You don’t know the question.”
“Historical experience.”
She ignored me.
A terminal in another region had developed what management called a safety-confidence score.
I thought I had misheard.
“A what?”
“Exactly.”
The idea sounded reasonable on paper.
Track how often individual drivers initiated safety holds.
Compare them with review outcomes.
Identify people who repeatedly reported situations later found unsupported.
Then provide coaching.
Not discipline.
Coaching.
I knew where this was going before Nina finished.
“They’re ranking drivers.”
“Unofficially.”
“Which means yes.”
“Yes.”
“How?”
“Green, yellow, red.”
I closed my eyes.
Red cells again.
“Who sees it?”
“Dispatch supervisors.”
“Drivers?”
“Not routinely.”
“Does it affect routes?”
“They say no.”
“That answer aged badly last time.”
“Exactly why I called.”
The problem was subtle.
If dispatch saw a red indicator beside a driver’s name, every new safety report started with doubt.
That could shape tone.
Escalation.
Trust.
Even if no explicit punishment occurred.
Nina sent me an anonymized example.
DRIVER A — 11 SAFETY HOLDS / 6 APPROVED / 5 INSUFFICIENT EVIDENCE.
Red.
I stared at it.
“What does insufficient evidence mean?”
“That’s one of the issues.”
“Could be legitimate concern that cleared before review.”
“Yes.”
“Could be weather that changed.”
“Yes.”
“Could be a mechanical noise nobody reproduced later.”
“Yes.”
“And they’re treating those like false reports?”
“Not officially.”
Again.
Officially was doing too much work.
I asked, “Who created this?”
“Local operations analyst.”
“Why?”
“To reduce unnecessary stops.”
There it was.
The new system had succeeded at protecting stops.
Now someone wanted efficiency back.
Nothing wrong with efficiency.
Until it quietly became a reason to distrust caution.
“Has anyone lost work?”
“Not documented.”
“Then what’s the harm so far?”
Nina went quiet.
Good question.
I had learned not to call something retaliation before facts supported it.
Finally she answered.
“Drivers found out.”
“How?”
“Screenshot.”
Of course.
“Reaction?”
“Some are already saying they won’t report marginal issues.”
That was the harm.
Not yet measurable in schedules.
Behavior changing in anticipation.
People silencing themselves because they feared a label.
“Kill the score.”
“You want to come say that?”
“No.”
“Coward.”
“Retired.”
Nina laughed.
The issue went before national safety.
Richard called me afterward.
Apparently I remained on some unofficial list of people executives contacted when they wanted a driver opinion without scheduling an entire panel.
“What’s your objection?” he asked.
“It measures outcomes as if uncertainty is failure.”
“Explain.”
“If I hear a brake noise and maintenance finds nothing, was the report wrong?”
“Not necessarily.”
“If I stop for wind and it drops ten minutes later?”
“Not necessarily.”
“If I report fatigue, rest thirty minutes, and feel better?”
“Not necessarily.”
“Then why color me red?”
Richard sighed.
“Because some drivers clearly overuse the process.”
“I believe that.”
“You do?”
“Yes.”
He sounded surprised again.
I was getting tired of that.
“Some people will abuse any protection.”
“Then how do we identify them?”
“Review behavior.”
“That’s what the score does.”
“No. The score replaces behavior with a number.”
Silence.
I continued.
“If someone has eleven stops, look at eleven stops.”
“That doesn’t scale.”
“Then build a better review system.”
“Expensive.”
“Crashes are expensive.”
“That argument can justify anything.”
Fair.
I thought again.
“What do you actually want to know?”
“Whether a driver is acting in good faith.”
“You can’t get good faith from a red dot.”
Richard was quiet.
That was the core problem.
Metrics were good at counting.
Bad at intention.
Still, he was right about scale.
A national company could not personally investigate every minor disagreement forever.
We needed something better than simply rejecting measurement.
I called Eric.
He had recently been promoted to senior dispatcher.
“You got five minutes?”
“Depends. Are you starting another policy revolution?”
“Maybe.”
He groaned.
I explained the score.
Eric hated it immediately.
“Dispatchers will trust the colors.”
“Even if told not to?”
“Especially then.”
“Why?”
“Because we’re busy.”
That honesty mattered.
“When you have thirty trucks moving and five problems at once, anything that tells you who’s ‘reliable’ becomes a shortcut.”
“So what would help without biasing you?”
He thought.
“Event flags, not driver flags.”
That was good.
“What kind?”
“If a specific safety report is unusual, flag the event for review.”
“Not the person.”
“Right.”
“How do you catch patterns?”
“Periodic independent audit.”
“Still expensive.”
“Less expensive than poisoning every call with a reputation score.”
Exactly.
Nina carried that proposal forward.
The national group eventually suspended the driver-level confidence scoring.
Instead, they built event-based review triggers.
Repeated unsupported reports could still be examined, but no red indicator appeared beside a driver’s name during live dispatch.
Historical patterns were visible only to designated safety reviewers, not operational dispatchers making immediate decisions.
Better separation.
Another firewall.
When Richard told me the decision, he said, “You know what worries me?”
“What?”
“We keep inventing ways to recreate the old problem with better software.”
I laughed.
“That’s management.”
“Comforting.”
He was right.
The danger had evolved.
Victor had used personal discretion.
Later supervisors used tone.
Now algorithms and dashboards could quietly encode suspicion.
Same pressure.
Different mechanism.
That became the next major theme in training.
Technology should support judgment without prejudging the person making it.
I attended one meeting as a guest.
Only one.
Nina threatened to lock the door if I tried leaving early.
The technology team demonstrated a prototype.
Weather overlays.
Vehicle fault data.
Hours-of-service status.
Traffic.
All visible during safety calls.
Impressive.
Then a product manager showed a proposed “risk confidence” percentage.
I raised my hand.
“No.”
He laughed.
“I haven’t explained it.”
“Still no.”
Nina smiled.
“Welcome to Marcus.”
The manager explained anyway.
The system would estimate whether available external data supported the driver’s reported hazard.
Wind.
Ice.
Mechanical alerts.
Traffic incidents.
“If confidence is low?” I asked.
“Dispatch sees that.”
“No.”
“Why?”
“Because absence of external confirmation is not evidence the driver is wrong.”
He frowned.
“It’s information.”
“Then show the information.”
“That’s what the score summarizes.”
“Summaries become conclusions.”
The room went quiet.
I pointed at the demo.
“If the weather service says no advisory, show that. If traffic cameras look clear, show that. If vehicle telemetry shows no fault, show that.”
“And let dispatch decide?”
“Yes.”
“That’s slower.”
“Yes.”
He looked frustrated.
I understood.
Software wanted neat outputs.
Human safety decisions were messy.
Rachel, attending remotely, spoke.
“Marcus’s concern mirrors what we found in the original incidents. Characterizations replaced facts.”
Exactly.
The confidence score disappeared from the final interface.
Facts remained visible.
No machine-generated verdict.
I felt oddly proud of that.
Not because technology was bad.
Because tools should illuminate reality, not quietly decide which human deserved belief.
Months passed.
The national safety program settled again.
Incidents still happened.
Complaints still appeared.
Some drivers were disciplined after reviews found obvious misuse.
One claimed unsafe weather while electronic records showed he had stopped at his house for nearly two hours.
Another repeatedly reported trailer defects that inspections found he had not actually checked before departure.
Those cases mattered.
A fair system had to protect credibility by addressing dishonesty too.
Nina worried the discipline would scare drivers.
It did, briefly.
But the company published anonymized explanations.
Facts.
Not names.
What was reported.
What evidence showed.
Why the decision was reached.
Transparency again.
Trust recovered.
At home, life changed in quieter ways.
Sarah and I refinanced the house.
The car finally stopped making the noise she had complained about for a year.
I replaced my work boots before the soles became smooth, which she treated as evidence of personal growth.
Ben’s daughter graduated high school.
Carla became a driver mentor.
Dennis retired.
His last shift ended with half the terminal standing outside pretending we had gathered accidentally.
He hated attention almost as much as I did.
Paula handed him a plaque.
He looked at me.
“Ugly?”
“Probably.”
He opened it.
“Definitely.”
Everyone laughed.
Dennis found me afterward.
“You still carrying the flashlight?”
“Yes.”
“Good.”
“Why?”
“Tradition.”
“You almost died because you ignored your own steering wheel and now you’re giving me superstition advice?”
“Retirement gives wisdom.”
“Apparently.”
He shook my hand.
Then said quietly, “Don’t let them forget the uncomfortable parts.”
I understood.
Once a program succeeded, organizations liked telling the cleaned-up version.
Problem discovered.
Policy fixed.
Metrics improved.
That version was easier.
Less useful.
The uncomfortable parts were where learning lived.
A warning removed only after an important customer intervened.
Drivers who stayed silent for years.
Records altered.
Good policies that still produced bad tone.
A crash under the new system.
Metrics that threatened to recreate bias.
None of that fit neatly into a success story.
It mattered anyway.
After Dennis retired, Paula asked whether I wanted his occasional trainer shifts.
I surprised both of us.
“Maybe.”
She stared.
“What?”
“Don’t make it weird.”
“You said maybe.”
“I can still change my mind.”
She smiled.
“We’ll start small.”
So I began training new drivers twice a month.
Mostly practical things.
Pre-trip inspections.
Backing.
Load paperwork.
Weather planning.
Electronic logs.
I refused motivational speeches.
On the first day, a new driver asked, “What’s the biggest mistake rookies make?”
I expected someone to say backing too fast.
Instead I answered, “Thinking uncertainty means they’re stupid.”
The room went quiet.
I explained.
A strange vibration.
Weather that feels worse than the forecast.
A trailer that doesn’t seem right.
A route you’re not sure you can safely finish.
“You’re allowed not to know immediately.”
One driver asked, “So just stop whenever?”
“No.”
“Then what?”
“Notice. Slow down. Gather information. Call. Make the safest reasonable decision you can with what you know.”
Another asked, “What if dispatch disagrees?”
“Then make them explain with facts.”
That became my training style.
Less certainty.
More process.
It seemed useful.
One evening after class, I found Lily’s flashlight on the desk.
I had brought it to show them emergency equipment.
A trainee named Alex picked it up.
“This thing looks ancient.”
“It’s not.”
“The tape is falling off.”
“Put it down.”
He laughed.
“What’s Lily?”
“Long story.”
He switched it on.
The beam flickered.
Then died.
I stared.
Alex tapped it.
“Battery?”
“Fresh.”
He opened the compartment.
Clean.
He tried again.
Nothing.
For years, that flashlight had survived rain, glove compartments, training rooms, roadside stops.
Now it simply stopped working.
I took it home.
Changed batteries again.
Cleaned the contacts.
Nothing.
Sarah found me at the kitchen table with a screwdriver.
“You know you can buy another flashlight.”
“It’s not the point.”
“I know.”
I tried one more time.
Dead.
Finally Sarah sat across from me.
“Maybe it’s done.”
I looked at the scratched plastic.
The faded tape.
LILY.
“I was supposed to keep it ready.”
“You did.”
That hurt in a stupid way.
I laughed at myself.
“It’s a flashlight.”
“Yes.”
“I’m upset about a flashlight.”
“Yes.”
She reached across the table.
“Because it isn’t just a flashlight.”
I hated when she was right.
The next day, I texted Julian.
Bad news. The official flashlight program has suffered equipment failure.
His response came quickly.
Lily says do not dispose of evidence.
Then:
She wants to inspect it.
A week later, an envelope arrived.
Inside was the old flashlight.
Repaired.
And another one.
Newer.
Brighter.
Taped to the new one was a note in Lily’s increasingly neat handwriting.
BACKUP SYSTEM.
I laughed so hard Sarah came from the other room.
She read it.
Then looked at me.
“Smart kid.”
“Terrifying.”
I placed both flashlights in my work bag.
One old.
One new.
Neither replacing the other.
That seemed right.
The company’s system had become something similar.
Old lessons preserved.
New tools added.
Redundancy.
Because eventually anything people depended on could fail.
The answer was not pretending failure would never happen.
It was building enough backup that one failure did not leave everyone in the dark.
Click here to continue reading: PART 23: Training New Drivers Forced Me to Tell the Storm Story Without Its Convenient Ending, and One Rookie Asked the Question I’d Avoided for Years
The Storm Was Already Costing Me Time When a Child’s Flashlight Appeared Through the Rain Beside the Highway
Part 22 of 30

