Repository navigation
OpenConceptLab/ocl_issues#2865 | rate limit 429 and capacity limit 429 labeling/logging - #93
Conversation
…9 labeling/logging
Codex review, pass 1 (codex-cli 0.161.0, commit 28f986d)Lean-loop defects pass, run at Verdict: findings: 1 finding(s).
findings.json{
"reviewer": "codex-cli 0.161.0",
"perspective": "defects",
"mode": "full",
"commit": "28f986d6ecd70fa7a11939622ae18aaaff0e096f",
"verdict": "findings",
"prior": [],
"findings": [
{
"id": "F1.1",
"severity": "P1",
"tag": null,
"file": "src/services/capacity.js",
"line": 233,
"claim": "Shared-gate waits display an incorrect retry time and can retain the wrong limit label.",
"scenario": "A request entering a 60-second rate-limit gate announces a retry at t+60s, but random()=0.8 adds 24 seconds of jitter afterward, so it actually sends at t+84s. If another 429 extends the gate or changes its limit, onWait is never refreshed: a reproduced capacity extension delayed sending until t+149s while the notice still said rate-limit retry at t+60s. The new UI derives its time and reason directly from this stale notification.",
"action": "Refresh wait metadata when the gate deadline or limit changes, and account for post-gate jitter in the displayed retry time without changing retry behavior. Add tests covering nonzero jitter and an extended gate whose limit changes.",
"lines_delta": null,
"confidence": "high"
}
]
} |
paynejd
left a comment
There was a problem hiding this comment.
Follow-up review (Claude), commit 28f986d
This follows Codex pass 1 above. The error_code split, the gate's limit, the i18n keys and the AC3 tests all look right to me. CI's lint and npm run test:all pass locally (412 unit, 6 integration). There are two things to fix before oclmap#86 goes out, plus a note.
1. Codex F1.1 holds (P1: the user only sees the wrong time; the retry behaviour itself is right). Most rows in a busy Auto Match run sit in the paused branch. That branch announces now + pauseMs, then sleeps up to another 50% of pauseMs once the gate opens (capacity.js#L236-L240). So the chip can show "retrying at 1:02:30" while the request is still waiting at 1:02:55. These lines are outside the diff hunk, so this can't be a one-click suggestion. Here is the patch instead (it passes the 412 tests and lint locally):
// Spread out the requests the pause held back, within the budget, and
// announce the retry time with that spread in it.
const jitterMs = Math.min(pauseMs * THROTTLE_JITTER * random(), budgetMs - pauseMs)
onWait?.({ reason: 'paused', delayMs: pauseMs + jitterMs, status: null, retryAfterMs: pauseMs, limit: gate.limit?.() ?? CAPACITY_LIMIT })
const startedAt = now()
let outcome = await gate.wait(isCancelled, { sleep, pollMs, maxMs: budgetMs })
if(outcome === true)
outcome = await sleepUnlessCancelled(Math.min(jitterMs, Math.max(budgetMs - (now() - startedAt), 0)), isCancelled, { sleep, pollMs })The patch doesn't cover the gate being extended (or changing limit) mid-wait. One way to cover it: wait on the gate only for the announced pauseMs. If it's still paused after that, call onWaitEnd and continue, so the loop's existing budget check runs and onWait fires again with the new time and limit.
2. A long rate limit still reads "Server busy" (P2, same mislabel as the ticket). oclapi2's match throttles include day windows. A day-limit 429's Retry-After can be hours, which exceeds both wait caps (5 min interactive, 30 min bulk). So requestWithCapacityRetry returns throttled at once and never calls onWait. The row, its log and the end-of-run notice then say "The server was busy… run it again later" (MapProject.jsx#L2121, #L2344, #L2585, #L3851, #L4150). A user who has hit the day limit is told to retry soon, and retrying fails again. The throttled result already carries the 429 in error.response, so getThrottleLimit can pick rate-limit wording ("limit reached, try again after …"). On the early exit before any attempt, lastError is undefined, so that path needs gate.limit() instead. This can go in this PR or in a sibling ticket, as long as it ships before oclmap#86.
3. Note (P3): the run log's capacity_wait description now includes a clock time in the browser's locale, with no date or timezone, in a stored log. extras already records limit and retry_after_ms, so the description could leave the time out. Also, pt-BR (oclmap#90) will need rate_limited and rate_limited_short if this merges first.
paynejd
left a comment
There was a problem hiding this comment.
Changes requested (Claude, commit 28f986d)
Two fixes are needed before this merges, and before oclmap#86 goes to production. Codex pass 1 and my earlier review have the detail.
- Codex F1.1: the retry time is wrong for requests held by the shared gate. The label shows a time up to 50% of the pause too early, and it doesn't update when another 429 extends the pause.
- The day-limit case: please fold it into this PR (Jon's call). A rate limit with a
Retry-Afterlonger than the wait cap gives up without waiting, and the row, its logs, the alert and the end-of-run notice all still say "server busy". With the fix, every throttled result carrieslimitandretryAt. When a rate limit was the refusal, the user sees "You've reached your request limit… run it again after {{time}}". Capacity refusals keep today's wording.
How to apply. Commit the 7 suggestions as one batch (Files changed → Add suggestion to batch → Commit suggestions). Then pull and run git apply with the patch below. Most of the change is outside this PR's diff hunks, so it can't be a suggestion. The suggestions on their own pass the tests and lint. With the patch, npm run test:all passes (418 unit, 6 integration), and so does CI's ESLint command.
The patch touches:
capacity.js: F1.1.matchBatch.js,MapProject.jsx,Candidates.jsx: carry{limit, retryAt}to the row label (newthrottleprop besidecapacityWait), thealgo_throttled/rerank_throttled/auto_match_rows_throttledlogs, the alert and the end-of-run notice. When a run mixes the two kinds, the rate limit with the latestretryAtwins the notice.rowProgress.js:row_rate_limitedfor a row a rate limit refused.- Six new tests.
Please read the MapProject.jsx part closely. It's the widest change, and I reviewed it by reading, not in the browser.
Still optional (P3): the capacity_wait log description still includes the clock time.
remaining.patch (apply after committing the suggestions)
diff --git a/src/components/map-projects/Candidates.jsx b/src/components/map-projects/Candidates.jsx
index cb2320e..7629ae6 100644
--- a/src/components/map-projects/Candidates.jsx
+++ b/src/components/map-projects/Candidates.jsx
@@ -377,7 +377,7 @@ const CandidateList = ({rowViews, header, rowIndex, sortBy, order, openConceptPa
// conceptCache — project-wide ConceptDefinition store, keyed by concept_key.
// algosSelected — algorithm definitions (for headers/grouping).
// (plans/unified-mapper-model.md "How the views map onto this model".)
-const Candidates = ({rowIndex, rowState, conceptCache, targetCanonical, targetRelativeUrl, openConceptPanel, showItem, isSelectedForMap, onMap, onFetchMore, isLoading, candidatesScore, repoVersion, analysis, onFetchRecommendation, appliedFacets, setAppliedFacets, filters, facets, columns, defaultFilters, locales, models, selectedModel, onModelChange, promptTemplates, promptTemplate, onPromptTemplateChange, onRefreshClick, rowStage, inAIAssistantGroup, algosSelected, canSelectAIModel, capacityWait=false}) => {
+const Candidates = ({rowIndex, rowState, conceptCache, targetCanonical, targetRelativeUrl, openConceptPanel, showItem, isSelectedForMap, onMap, onFetchMore, isLoading, candidatesScore, repoVersion, analysis, onFetchRecommendation, appliedFacets, setAppliedFacets, filters, facets, columns, defaultFilters, locales, models, selectedModel, onModelChange, promptTemplates, promptTemplate, onPromptTemplateChange, onRefreshClick, rowStage, inAIAssistantGroup, algosSelected, canSelectAIModel, capacityWait=false, throttle}) => {
const { t } = useTranslation();
const [sortBy, setSortBy] = React.useState('rerank_score')
const [groupBy, setGroupBy] = React.useState('quality')
@@ -436,7 +436,7 @@ const Candidates = ({rowIndex, rowState, conceptCache, targetCanonical, targetRe
const canFetchMore = hasAnyView
const algoStagesValue = values(rowStage || {}).filter((_, i) => Object.keys(rowStage || {})[i] !== 'recommend')
const areAlgoRun = algoStagesValue.length > 0 && algoStagesValue.every(v => v === 1)
- const { label, status: progressStatus } = getRowProgressLabel(rowStage, algosSelected, {t, capacityWait});
+ const { label, status: progressStatus } = getRowProgressLabel(rowStage, algosSelected, {t, capacityWait, throttle})
// Waiting out a busy server, or given up on for now: not a spinner (ocl_issues#2849).
const isCapacityStatus = ['capacity_wait', 'throttled'].includes(progressStatus)
diff --git a/src/components/map-projects/MapProject.jsx b/src/components/map-projects/MapProject.jsx
index 74425f1..7564b9c 100644
--- a/src/components/map-projects/MapProject.jsx
+++ b/src/components/map-projects/MapProject.jsx
@@ -71,7 +71,7 @@ import { OperationsContext } from '../app/LayoutContext';
import APIService, { isTransientNetworkError, retryWithBackoff } from '../../services/APIService';
import { buildAttributionHeaders, buildConfigSnapshot, summarizeRunCompletion, matchesOnOCL } from '../../services/attribution'
-import { CAPACITY_WAIT_CAP_MS, HEAVY_REQUEST_TIMEOUT_MS, LIGHT_REQUEST_TIMEOUT_MS, createCapacityGate, createLimiter, isRetryableError, requestWithCapacityRetry, sleepUnlessCancelled, untilCancelled } from '../../services/capacity'
+import { CAPACITY_LIMIT, RATE_LIMIT, CAPACITY_WAIT_CAP_MS, HEAVY_REQUEST_TIMEOUT_MS, LIGHT_REQUEST_TIMEOUT_MS, createCapacityGate, createLimiter, isRetryableError, requestWithCapacityRetry, sleepUnlessCancelled, untilCancelled } from '../../services/capacity'
import { highlightTexts, dropVersion, getCurrentUser, hasAuthGroup, hasCapability, getMapperPreview, getNewProjectBlockReason, downloadObject, currentUserToken, refreshCurrentUserCapabilitiesCache } from '../../common/utils';
import { WHITE, SURFACE_COLORS, TEXT_GRAY } from '../../common/colors';
@@ -116,7 +116,7 @@ import { normalizeAlgorithmInvocation, hasSuccessfulAlgorithmResponse, getAlgori
import { parseConceptKey } from './conceptKey'
import { getDefaultTargetRepoVersion, getProjectTargetRepoVersion, getTargetRepoVersionFromUrl, getTargetRepoVersionId } from './projectTargetRepo'
import { buildBridgeTargetDownloadEntries, buildQualityRowViews, conceptBelongsToTargetRepo, conceptForMapping, formatBridgeTargetDownloadEntry, resolveAICandidateID, getScoreDetails, getAIAnalysisCandidateIDs } from './viewBuilders.js'
-import { getCapacityWaitLabel, mergeCapacityWaits } from './rowProgress.js'
+import { formatClockTime, getCapacityWaitLabel, mergeCapacityWaits } from './rowProgress.js'
import './MapProject.scss'
import '../common/ResizablePanel.scss'
@@ -245,6 +245,7 @@ const MapProject = () => {
// A run's rows the server stayed too busy for, by algorithm id ('rerank'
// included), for the end-of-run notice.
const throttledRunRowsRef = React.useRef({})
+ const rowThrottlesRef = React.useRef({})
// One logs POST at a time per project URL (see createLatestSender).
const logsSendersRef = React.useRef(new Map())
// The row a bridge $match is for, set around the call into the Bridge Match
@@ -1342,6 +1343,7 @@ const MapProject = () => {
setAlert(false)
setSelectedCandidatesScoreBucket(false)
setScoreBucketSortBy('desc')
+ rowThrottlesRef.current = {}
setRowStage({})
}
@@ -1820,21 +1822,34 @@ const MapProject = () => {
scheduleAutoSave('auto_match_stopped')
}
- const noteThrottledRunRow = (algoId, index) => {
+ // Refusals retain the latest rate limit, even without a wait (ocl_issues#2865).
+ const mergeThrottles = (a, b) => a?.limit === RATE_LIMIT && (b?.limit !== RATE_LIMIT || a.retryAt > b.retryAt) ? a : b
+ const getThrottleExtras = throttle => ({
+ limit: throttle?.limit || CAPACITY_LIMIT,
+ ...(throttle?.limit === RATE_LIMIT ? {retry_at: throttle.retryAt} : {}),
+ })
+ const getThrottledRowDescription = throttle => t(throttle?.limit === RATE_LIMIT ? 'map_project.row_rate_limited_log' : 'map_project.row_throttled')
+ const getThrottledAlgorithmMessage = throttle => throttle?.limit === RATE_LIMIT ?
+ t('map_project.algorithm_rate_limited', {time: formatClockTime(throttle.retryAt)}) : t('map_project.algorithm_throttled')
+
+ const noteThrottledRunRow = (algoId, index, throttle) => {
throttledRunRowsRef.current = {
...throttledRunRowsRef.current,
- [algoId]: uniq([...(throttledRunRowsRef.current[algoId] || []), index]),
+ [algoId]: {
+ ...mergeThrottles(throttledRunRowsRef.current[algoId], {limit: throttle?.limit, retryAt: throttle?.retryAt}),
+ rows: uniq([...(throttledRunRowsRef.current[algoId]?.rows || []), index]),
+ },
}
}
// Once a request for an algorithm (or rerank) stayed refused for the whole
// cap, the run stops asking for it: the rest of its rows end throttled at
// once, "not run, retry", instead of each waiting out the cap again.
- const isRunThrottled = algoId => Boolean(throttledRunRowsRef.current[algoId]?.length)
- const markRowsThrottled = (indexes, algoId, logExtras = {}) => indexes.forEach(index => {
- markAlgo(index, algoId, -4)
- log({action: algoId === 'rerank' ? 'rerank_throttled' : 'algo_throttled', description: t('map_project.row_throttled'), extras: {...logExtras, skipped: true}}, index)
- noteThrottledRunRow(algoId, index)
+ const isRunThrottled = algoId => Boolean(throttledRunRowsRef.current[algoId]?.rows.length)
+ const markRowsThrottled = (indexes, algoId, logExtras = {}, throttle = throttledRunRowsRef.current[algoId]) => indexes.forEach(index => {
+ markAlgo(index, algoId, -4, throttle)
+ log({action: algoId === 'rerank' ? 'rerank_throttled' : 'algo_throttled', description: getThrottledRowDescription(throttle), extras: {...logExtras, ...getThrottleExtras(throttle), skipped: true}}, index)
+ noteThrottledRunRow(algoId, index, throttle)
})
// ── ocl_online#105 Phase 5: AutomatchRun attribution ──────────────────────
@@ -2118,7 +2133,7 @@ const MapProject = () => {
if(result.ok)
return result.response
if(result.reason === 'throttled')
- return {detail: t('map_project.algorithm_throttled'), status: 429, throttled: true}
+ return {detail: getThrottledAlgorithmMessage(result), status: 429, throttled: true, limit: result.limit, retryAt: result.retryAt}
if(result.reason === 'cancelled')
return {detail: 'cancelled', cancelled: true}
const data = result.error?.response?.data
@@ -2329,9 +2344,9 @@ const MapProject = () => {
retryOptions: {gate: getCapacityGate(service.URL), onCapacity: noteCapacity},
...trackCapacityWait(rowIndexes, logExtras),
// A batch that outlives its run must not touch the next run's stages.
- setStage: (index, stage) => {
+ setStage: (index, stage, throttle) => {
if(isCurrentRun())
- markAlgo(index, algo.id, stage)
+ markAlgo(index, algo.id, stage, throttle)
},
onRowFinished: index => log({action: 'algo_finished', extras: logExtras}, index),
onRowFailed: (index, {error, status, attempts, previewLimit}) => {
@@ -2340,10 +2355,10 @@ const MapProject = () => {
failedMatchRows[algo.id] = [...(failedMatchRows[algo.id] || []), index]
},
// Still refused after the long cap: not run, and not a failure.
- onRowThrottled: (index, {attempts, waitedMs}) => {
- log({action: 'algo_throttled', description: t('map_project.row_throttled'), extras: {...logExtras, attempts, waited_ms: waitedMs}}, index)
+ onRowThrottled: (index, result) => {
+ log({action: 'algo_throttled', description: getThrottledRowDescription(result), extras: {...logExtras, ...getThrottleExtras(result), attempts: result.attempts, waited_ms: result.waitedMs}}, index)
if(isCurrentRun())
- noteThrottledRunRow(algo.id, index)
+ noteThrottledRunRow(algo.id, index, result)
},
// Sets matchQuotaStopRef, which stops the rest of the $match requests
// but lets the run finish with what it has.
@@ -2504,6 +2519,7 @@ const MapProject = () => {
// Reset all algo stages to -1 for every row before starting so that
// stages from a previous run don't make the "all algos done" check
// pass prematurely when only the first algo has finished.
+ rowsToProcess.forEach(row => { delete rowThrottlesRef.current[row.__index] })
setRowStage(prev => {
const next = { ...prev }
rowsToProcess.forEach(row => {
@@ -2580,11 +2596,17 @@ const MapProject = () => {
// Rows the server stayed too busy for weren't run: say so apart from
// the failures (ocl_issues#2849).
if(!isRunStopped() && keys(throttledRunRowsRef.current).length) {
- const rowsByAlgo = getRowsByAlgo(throttledRunRowsRef.current)
- projectLog({action: 'auto_match_rows_throttled', extras: {row_indexes_by_algorithm: rowsByAlgo}})
- endOfRunNotices.push(t('map_project.auto_match_rows_throttled', {
- count: uniq(flatten(values(rowsByAlgo))).length,
- details: describeRows(rowsByAlgo),
+ const rowsByAlgo = getRowsByAlgo(Object.fromEntries(Object.entries(throttledRunRowsRef.current).map(([algoId, info]) => [algoId, info.rows])))
+ const throttle = values(throttledRunRowsRef.current).reduce(mergeThrottles, null)
+ const rateLimited = throttle.limit === RATE_LIMIT
+ const details = {count: uniq(flatten(values(rowsByAlgo))).length, details: describeRows(rowsByAlgo)}
+ projectLog({
+ action: 'auto_match_rows_throttled',
+ description: t(rateLimited ? 'map_project.auto_match_rows_rate_limited_log' : 'map_project.auto_match_rows_throttled', details),
+ extras: {row_indexes_by_algorithm: rowsByAlgo, ...getThrottleExtras(throttle)},
+ })
+ endOfRunNotices.push(t(rateLimited ? 'map_project.auto_match_rows_rate_limited' : 'map_project.auto_match_rows_throttled', {
+ ...details, ...(rateLimited ? {time: formatClockTime(throttle.retryAt)} : {}),
}))
}
if(endOfRunNotices.length)
@@ -2739,7 +2761,7 @@ const MapProject = () => {
};
if(isRunThrottled(algo.id) || isRunThrottled('ocl-scispacy-loinc')) {
- markRowsThrottled(map(_rows.slice(index), '__index'), algo.id, getAlgoLogExtras(algo))
+ markRowsThrottled(map(_rows.slice(index), '__index'), algo.id, getAlgoLogExtras(algo), throttledRunRowsRef.current[isRunThrottled(algo.id) ? algo.id : 'ocl-scispacy-loinc'])
break
}
markAlgo(_rows[index].__index, algo.id, 0)
@@ -3714,7 +3736,11 @@ const MapProject = () => {
return next;
};
- const markAlgo = (rowId, algoId, value) => {
+ const markAlgo = (rowId, algoId, value, throttle) => {
+ if(value === -4 && throttle)
+ rowThrottlesRef.current[rowId] = mergeThrottles(rowThrottlesRef.current[rowId], {limit: throttle.limit, retryAt: throttle.retryAt})
+ else if(value === 0 && !values(omit(rowStageRef.current[rowId], algoId)).includes(-4))
+ delete rowThrottlesRef.current[rowId]
setRowStage(prev => {
const needsRerank = isMultiAlgo || find(algosSelected, { type: "custom" }) || find(algosSelected, { type: "ocl-scispacy" });
const row = ensureRow(prev, rowId, selectedAlgoIds, needsRerank);
@@ -3844,11 +3870,11 @@ const MapProject = () => {
gate: getCapacityGate(service.URL),
onCapacity: noteCapacity,
...trackCapacityWait([__row.__index], getAlgoLogExtras(algoDef)),
- }}).then(({ok, response, errorBody, previewLimit, throttled}) => {
+ }}).then(({ok, response, errorBody, previewLimit, throttled, limit, retryAt}) => {
if(ok)
return callback(response, payload)
if(throttled)
- return callback({detail: t('map_project.algorithm_throttled'), status: 429, throttled: true}, payload)
+ return callback({detail: getThrottledAlgorithmMessage({limit, retryAt}), status: 429, throttled: true, limit, retryAt}, payload)
// A preview limit keeps the server's body, which opens the limit dialog.
callback(previewLimit ? errorBody : {...errorBody, detail: t('map_project.match_request_failed', {error: errorBody.detail})}, payload)
})
@@ -3871,6 +3897,8 @@ const MapProject = () => {
if(!algoId)
return
let __row = isEmpty(_row) ? row : _row
+ if(!keepAlert)
+ delete rowThrottlesRef.current[__row.__index]
// Reuse when the algo has already produced an AlgorithmResponse for
// this row, regardless of whether it returned matches. Gating on
@@ -3936,15 +3964,15 @@ const MapProject = () => {
// The server stayed too busy: the algorithm didn't run, it didn't
// fail, and no failed response is kept (ocl_issues#2849).
const isThrottled = Boolean(response.throttled)
- log({action: isThrottled ? 'algo_throttled' : 'algo_failed', ...(isThrottled ? {description: t('map_project.row_throttled')} : {}), extras: {...logExtras, error: response.detail, status: response.status, ...(offset ? {offset} : {})}}, __row.__index)
+ log({action: isThrottled ? 'algo_throttled' : 'algo_failed', ...(isThrottled ? {description: getThrottledRowDescription(response)} : {}), extras: {...logExtras, ...(isThrottled ? getThrottleExtras(response) : {}), error: isThrottled ? getThrottledRowDescription(response) : response.detail, status: response.status, ...(offset ? {offset} : {})}}, __row.__index)
setAlert({message: response.detail, severity: isThrottled ? 'warning' : 'error'})
if(offset) {
// A failed "load more" keeps the pages already loaded. The
// algorithm stays done if it loaded a page; "load more" also runs
// algorithms that never did, and those stay failed.
- markAlgo(__row.__index, algoId, hasSuccessfulAlgorithmResponse(rowMatchStateRef.current?.[__row.__index], algoId) ? 1 : (isThrottled ? -4 : -2))
+ markAlgo(__row.__index, algoId, hasSuccessfulAlgorithmResponse(rowMatchStateRef.current?.[__row.__index], algoId) ? 1 : (isThrottled ? -4 : -2), response)
} else if(isThrottled) {
- markAlgo(__row.__index, algoId, -4)
+ markAlgo(__row.__index, algoId, -4, response)
} else {
markAlgo(__row.__index, algoId, -2)
mergeIntoRowMatchState(__row.__index, normalizeAlgorithmInvocation(null, {
@@ -4146,11 +4174,11 @@ const MapProject = () => {
if(isSuperseded())
return
// Not run, not failed: the row can be run again.
- markAlgo(__row.__index, SCISPACY_ALGO_ID, -4)
- log({action: 'algo_throttled', description: t('map_project.row_throttled'), extras: {algo: SCISPACY_ALGO_ID, attempts: result.attempts, waited_ms: result.waitedMs}}, __row.__index)
+ markAlgo(__row.__index, SCISPACY_ALGO_ID, -4, result)
+ log({action: 'algo_throttled', description: getThrottledRowDescription(result), extras: {algo: SCISPACY_ALGO_ID, ...getThrottleExtras(result), attempts: result.attempts, waited_ms: result.waitedMs}}, __row.__index)
if(isBulk)
- noteThrottledRunRow(SCISPACY_ALGO_ID, __row.__index)
- setAlert(prev => (prev?.severity === 'error' ? prev : {message: t('map_project.algorithm_throttled'), severity: 'warning'}))
+ noteThrottledRunRow(SCISPACY_ALGO_ID, __row.__index, result)
+ setAlert(prev => (prev?.severity === 'error' ? prev : {message: getThrottledAlgorithmMessage(result), severity: 'warning'}))
setIsLoadingInDecisionView(false)
onFailure?.()
return
@@ -4394,12 +4422,12 @@ const MapProject = () => {
return null
}
if(result.reason === 'throttled') {
- log({action: 'rerank_throttled', description: t('map_project.row_throttled'), extras: {attempts: result.attempts, waited_ms: result.waitedMs}}, index)
+ log({action: 'rerank_throttled', description: getThrottledRowDescription(result), extras: {...getThrottleExtras(result), attempts: result.attempts, waited_ms: result.waitedMs}}, index)
if(isSuperseded())
return null
- markAlgo(index, 'rerank', -4)
+ markAlgo(index, 'rerank', -4, result)
if(isRunTraffic)
- noteThrottledRunRow('rerank', index)
+ noteThrottledRunRow('rerank', index, result)
return null
}
if(!result.ok) {
@@ -4753,10 +4781,10 @@ const MapProject = () => {
// The server stayed too busy: not run, not failed.
if(response?.throttled) {
clearRefreshRowStageSnapshot(__row.__index)
- markAlgo(__row.__index, bridgeAlgoId, -4)
- log({action: 'algo_throttled', description: t('map_project.row_throttled'), extras: getAlgoLogExtras(bridgeAlgo)}, __row.__index)
+ markAlgo(__row.__index, bridgeAlgoId, -4, response)
+ log({action: 'algo_throttled', description: getThrottledRowDescription(response), extras: {...getAlgoLogExtras(bridgeAlgo), ...getThrottleExtras(response)}}, __row.__index)
if(isBulk)
- noteThrottledRunRow(bridgeAlgoId, __row.__index)
+ noteThrottledRunRow(bridgeAlgoId, __row.__index, response)
setAlert(prev => (prev?.severity === 'error' ? prev : {message: response.detail, severity: 'warning'}))
setIsLoadingInDecisionView(false)
onFailure?.()
@@ -6475,6 +6503,7 @@ const MapProject = () => {
rowIndex={rowIndex}
rowStage={rowStageRef.current[rowIndex]}
capacityWait={capacityWaits?.rows?.[rowIndex] || false}
+ throttle={rowThrottlesRef.current[rowIndex]}
rowState={rowMatchStateRef.current[rowIndex]}
conceptCache={conceptCache}
targetCanonical={buildProjectContext()?.target_repo?.canonical_url}
diff --git a/src/components/map-projects/__tests__/rowProgress.test.js b/src/components/map-projects/__tests__/rowProgress.test.js
index 4775c97..b539e91 100644
--- a/src/components/map-projects/__tests__/rowProgress.test.js
+++ b/src/components/map-projects/__tests__/rowProgress.test.js
@@ -10,7 +10,7 @@
import test from 'node:test'
import assert from 'node:assert/strict'
-import { getCapacityWaitLabel, getRowProgressLabel, mergeCapacityWaits } from '../rowProgress.js'
+import { formatClockTime, getCapacityWaitLabel, getRowProgressLabel, mergeCapacityWaits } from '../rowProgress.js'
import { CAPACITY_LIMIT, RATE_LIMIT } from '../../../services/capacity.js'
const ALGOS = [{id: 'ocl-semantic'}, {id: 'ocl-bridge'}]
@@ -111,3 +111,31 @@ test('mergeCapacityWaits: a capacity wait wins; between rate-limit waits, the la
assert.equal(mergeCapacityWaits(rateLimitWait, later), later)
assert.equal(mergeCapacityWaits(later, rateLimitWait), later)
})
+
+// Refused rows retain the limit even when no wait was announced (ocl_issues#2865).
+test('getRowProgressLabel: a refused rate limit shows when the row or rerank can run again', () => {
+ for(const stages of [
+ {'ocl-semantic': -4, 'ocl-bridge': 1},
+ {'ocl-semantic': 1, 'ocl-bridge': 1, rerank: -4},
+ ]) {
+ assert.deepEqual(getRowProgressLabel(stages, ALGOS, {t: tWith, throttle: rateLimitWait, formatTime}), {
+ label: 'map_project.row_rate_limited {"time":"t+20000"}', status: 'throttled',
+ })
+ assert.deepEqual(getRowProgressLabel(stages, ALGOS, {t: tWith, throttle: capacityWait, formatTime}), {
+ label: 'map_project.row_throttled', status: 'throttled',
+ })
+ }
+})
+
+test('formatClockTime: a retry today is a clock time; another day includes the weekday', () => {
+ const now = () => new Date(2026, 9, 7, 12).getTime()
+ const format = (date, options) => ({day: date.getDate(), ...options})
+ assert.deepEqual(formatClockTime(new Date(2026, 9, 7, 15).getTime(), {now, format}), {
+ day: 7, hour: 'numeric', minute: '2-digit', second: '2-digit',
+ })
+ for(const [date, day] of [[new Date(2026, 9, 8, 15), 8], [new Date(2026, 10, 7, 15), 7]]) {
+ assert.deepEqual(formatClockTime(date.getTime(), {now, format}), {
+ day, hour: 'numeric', minute: '2-digit', second: '2-digit', weekday: 'short',
+ })
+ }
+})
diff --git a/src/components/map-projects/matchBatch.js b/src/components/map-projects/matchBatch.js
index 1352f68..971dd7e 100644
--- a/src/components/map-projects/matchBatch.js
+++ b/src/components/map-projects/matchBatch.js
@@ -48,10 +48,10 @@ const getFailure = (err, attempts) => ({
* @param {object} opts
* @param {number[]} opts.rowIndexes the batch's row __index values
* @param {function} opts.send attempt => Promise<axios response>; must reject on failure
- * @param {function} opts.setStage (rowIndex, stage) => void
+ * @param {function} opts.setStage (rowIndex, stage, throttle) => void
* @param {function} opts.onRowFinished rowIndex => void
* @param {function} opts.onRowFailed (rowIndex, {error, status, attempts, previewLimit}) => void
- * @param {function} [opts.onRowThrottled] (rowIndex, {attempts, waitedMs}) => void
+ * @param {function} [opts.onRowThrottled] (rowIndex, {attempts, waitedMs, limit, retryAt}) => void
* @param {function} [opts.onPreviewLimit] err => void, for a preview-limit 403 (never retried)
* @param {function} [opts.onWait] info => void, as each wait starts (see requestWithCapacityRetry)
* @param {function} [opts.onWaitEnd] () => void, as each wait ends
@@ -84,8 +84,8 @@ export const runMatchBatch = async ({
}
if(result.reason === 'throttled') {
rowIndexes.forEach(index => {
- setStage(index, -4)
- onRowThrottled?.(index, {attempts: result.attempts, waitedMs: result.waitedMs})
+ setStage(index, -4, {limit: result.limit, retryAt: result.retryAt})
+ onRowThrottled?.(index, {attempts: result.attempts, waitedMs: result.waitedMs, limit: result.limit, retryAt: result.retryAt})
})
return []
}
@@ -199,5 +199,6 @@ export const requestSingleMatch = async (send, { retryOptions = {} } = {}) => {
errorBody: hasServerBody ? data : { detail: error, status },
previewLimit: isPreviewLimitError(err),
throttled: result.reason === 'throttled',
+ ...(result.reason === 'throttled' ? {limit: result.limit, retryAt: result.retryAt} : {}),
}
}
diff --git a/src/components/map-projects/rowProgress.js b/src/components/map-projects/rowProgress.js
index 730966c..725d743 100644
--- a/src/components/map-projects/rowProgress.js
+++ b/src/components/map-projects/rowProgress.js
@@ -39,13 +39,15 @@ export const mergeCapacityWaits = (a, b) => {
* server stayed too busy for (-4) asks for a retry: it wasn't run, it didn't
* fail.
*/
-export const getRowProgressLabel = (stageMap, algos, { t, capacityWait = false, formatTime } = {}) => {
+export const getRowProgressLabel = (stageMap, algos, { t, capacityWait = false, throttle, formatTime = formatClockTime } = {}) => {
if(stageMap === undefined)
return {label: false}
if(capacityWait)
return {label: getCapacityWaitLabel(capacityWait, {t, formatTime}), status: 'capacity_wait'}
if(!stageMap)
return {label: 'Preparing...', status: 'partial'}
+ const throttledLabel = () => throttle?.limit === RATE_LIMIT ?
+ t('map_project.row_rate_limited', {time: formatTime(throttle.retryAt)}) : t('map_project.row_throttled')
const stages = algos.map(k => stageMap[k.id]);
@@ -65,7 +67,7 @@ export const getRowProgressLabel = (stageMap, algos, { t, capacityWait = false,
if (stages.every(v => v === 1 || v === -3)) {
// The candidates are in, but the server stayed too busy to rank them.
if(stageMap.rerank === -4)
- return {label: t('map_project.row_throttled'), status: 'throttled'}
+ return {label: throttledLabel(), status: 'throttled'}
return true
}
@@ -75,7 +77,7 @@ export const getRowProgressLabel = (stageMap, algos, { t, capacityWait = false,
}
if(stages.some(v => v === -4))
- return {label: t('map_project.row_throttled'), status: 'throttled'}
+ return {label: throttledLabel(), status: 'throttled'}
return { label: 'Partially completed', status: 'partial' };
}
diff --git a/src/services/__tests__/capacity.test.js b/src/services/__tests__/capacity.test.js
index 99a4641..610647b 100644
--- a/src/services/__tests__/capacity.test.js
+++ b/src/services/__tests__/capacity.test.js
@@ -650,3 +650,69 @@ test('requestWithCapacityRetry: an error backoff carries no limit', async () =>
})
assert.deepEqual(waits, [['error', null]])
})
+
+// Retry labels and refusals keep their time and limit (ocl_issues#2865).
+test('requestWithCapacityRetry: a paused wait announces pause plus jitter, clamped to the budget', async () => {
+ for(const [maxWaitMs, expected] of [[20000, 12500], [11000, 11000]]) {
+ const clock = virtualClock()
+ const gate = createCapacityGate({now: clock.now})
+ gate.pause(10000, RATE_LIMIT)
+ const waits = []
+ const result = await requestWithCapacityRetry(scripted(ok()).send, {
+ gate, maxWaitMs, now: clock.now, sleep: clock.sleep, random: () => 0.5,
+ onWait: info => waits.push(info.delayMs),
+ })
+ assert.equal(result.ok, true)
+ assert.deepEqual(waits, [expected])
+ assert.equal(result.waitedMs, expected)
+ }
+})
+
+test('requestWithCapacityRetry: an extended gate re-announces its new limit and retry time', async () => {
+ const clock = virtualClock()
+ const gate = createCapacityGate({now: clock.now})
+ gate.pause(1000, CAPACITY_LIMIT)
+ const events = []
+ const result = await requestWithCapacityRetry(scripted(ok()).send, {
+ gate, now: clock.now, random: () => 0.5,
+ sleep: async ms => {
+ await clock.sleep(ms)
+ if(clock.t === 250)
+ gate.pause(2750, RATE_LIMIT)
+ },
+ onWait: info => events.push([info.limit, clock.t + info.delayMs]),
+ onWaitEnd: () => events.push('end'),
+ })
+ assert.equal(result.ok, true)
+ assert.deepEqual(events, [[CAPACITY_LIMIT, 1250], 'end', [RATE_LIMIT, 3500], 'end'])
+ assert.equal(result.waitedMs, 3500)
+})
+
+test('requestWithCapacityRetry: a long rate limit returns the server retry time without waiting', async () => {
+ const clock = virtualClock()
+ clock.t = 10000
+ const waits = []
+ const result = await requestWithCapacityRetry(scripted(throttled(86400)).send, {
+ now: clock.now, sleep: clock.sleep, random: () => 0.5, onWait: info => waits.push(info),
+ })
+ assert.equal(result.reason, 'throttled')
+ assert.equal(result.limit, RATE_LIMIT)
+ assert.equal(result.retryAt, 86410000)
+ assert.equal(result.waitedMs, 0)
+ assert.deepEqual(waits, [])
+})
+
+test('requestWithCapacityRetry: a gate refusal before any attempt carries its limit and retry time', async () => {
+ for(const limit of [CAPACITY_LIMIT, RATE_LIMIT]) {
+ const clock = virtualClock()
+ clock.t = 10000
+ const gate = createCapacityGate({now: clock.now})
+ gate.pause(60000, limit)
+ const {send, sent} = scripted(ok())
+ const result = await requestWithCapacityRetry(send, {gate, maxWaitMs: 1000, now: clock.now, sleep: clock.sleep})
+ assert.equal(result.reason, 'throttled')
+ assert.equal(result.limit, limit)
+ assert.equal(result.retryAt, 70000)
+ assert.deepEqual(sent, [])
+ }
+})
diff --git a/src/services/capacity.js b/src/services/capacity.js
index 51229aa..7cf5dbd 100644
--- a/src/services/capacity.js
+++ b/src/services/capacity.js
@@ -228,22 +228,26 @@ export const requestWithCapacityRetry = async (send, {
if(gate?.isPaused()) {
const pauseMs = gate.pausedForMs()
const budgetMs = maxWaitMs - waitedMs
+ const limit = gate.limit?.() ?? CAPACITY_LIMIT
if(pauseMs > budgetMs)
- return end({ ok: false, reason: 'throttled', error: lastError, limit: gate.limit?.() ?? CAPACITY_LIMIT, retryAt: now() + pauseMs })
- onWait?.({ reason: 'paused', delayMs: pauseMs, status: null, retryAfterMs: pauseMs, limit: gate.limit?.() ?? CAPACITY_LIMIT })
+ return end({ ok: false, reason: 'throttled', error: lastError, limit, retryAt: now() + pauseMs })
+ // Announce the jitter too, and refresh an extended pause (ocl_issues#2865).
+ const jitterMs = Math.min(pauseMs * THROTTLE_JITTER * random(), budgetMs - pauseMs)
+ onWait?.({ reason: 'paused', delayMs: pauseMs + jitterMs, status: null, retryAfterMs: pauseMs, limit })
const startedAt = now()
- let outcome = await gate.wait(isCancelled, { sleep, pollMs, maxMs: budgetMs })
+ let outcome = await gate.wait(isCancelled, { sleep, pollMs, maxMs: pauseMs })
// Spread out the requests the pause held back, within the budget.
if(outcome === true) {
- const jitterMs = Math.min(pauseMs * THROTTLE_JITTER * random(), Math.max(budgetMs - (now() - startedAt), 0))
- outcome = await sleepUnlessCancelled(jitterMs, isCancelled, { sleep, pollMs })
+ outcome = await sleepUnlessCancelled(Math.min(jitterMs, Math.max(budgetMs - (now() - startedAt), 0)), isCancelled, { sleep, pollMs })
}
waitedMs += now() - startedAt
onWaitEnd?.()
if(outcome === false)
return cancelled()
- if(outcome === 'timeout' || waitedMs > maxWaitMs)
- return end({ ok: false, reason: 'throttled', error: lastError })
+ if(outcome === 'timeout')
+ continue
+ if(waitedMs > maxWaitMs)
+ return end({ ok: false, reason: 'throttled', error: lastError, limit: gate.limit?.() ?? limit, retryAt: now() + gate.pausedForMs() })
continue
}
Co-authored-by: Jonathan Payne <paynejd@gmail.com>
…(capacity and rate-limit notices, Auto Match blocked reasons, search, pricing link) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… (input language, AI locales, ScispaCy preview limit, match metrics) Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Linked Issue
Closes OpenConceptLab/ocl_issues#2865