DEV Community

Asura
Asura

Posted on AI-assisted

I have created 2 formulas for AI/DL/ML. But i want someone to discover and check them and give response. This is my first time if their is wrong then sorry

============================================================================
FORMULA IDENTIFIER : AGS-Sci-v4.2.4-GRADIENT-FLOW-SCALER

MATHEMATICAL TYPE : Single-Variable Bounded Hyperbolic Activation Multiplier

FORMULA:
S(g) = 0.00407745501988 + 0.973576786100 * tanh(g)

DEFINITION OF VARIABLES:
g : Input Gradient Magnitude (The pre-activation backpropagated tensor signal)
S : Optimal Layer Update Scaling Coefficient (The resulting stabilizing multiplier)

HOW IT WORKS:
During the backward pass of a deep neural network, gradient signals often
suffer from vanishing or exploding values, disrupting parameter optimization.
This formula acts as an automated gating valve. As the gradient magnitude (g)
grows massive, the tanh(g) operator smoothly caps out at 1.0. This limits
the multiplier to approximately 0.977, applying a gentle mathematical brake
to prevent optimization spikes while ensuring a minimum update scale of
0.004077 is preserved for tiny gradients.

HOW TO USE IT (NUMPY BACKPROPAGATION NODE):

import numpy as np

def ags_gradient_scaler_backward(incoming_gradient):
# Compute the symmetrical scaling multiplier vector using absolute magnitude
S_g = 0.00407745501988 + 0.973576786100 * np.tanh(np.abs(incoming_gradient))

# Return the stabilized gradient back to the preceding layer
return incoming_gradient * S_g
Enter fullscreen mode Exit fullscreen mode

And

FORMULA IDENTIFIER : AGS-Sci-v4.2.4-ATTENTION-GATE-LAW

MATHEMATICAL TYPE : Multi-Variable Coupled Transcendental Memory Scaling Gate

FORMULA:
A(q, v) = tanh(0.312754428*q + 0.118784360*q*v + 0.091322486*q^2 + 0.040562366*v)

DEFINITION OF VARIABLES:
q : Normalized Query-Key Inner Product Vector (q = QK^T / sqrt(d_k))
v : Variance Footprint of the Context Window Matrix Sequence
A : Optimal Attention Gating Coefficient (Replaces standard Softmax mapping)

HOW IT WORKS:
Standard Softmax attention forces extreme focus on single tokens, often
leading to context drift or high-dimensional memory allocation collapse.
This formula wraps a complex cross-variable polynomial inside a global tanh
blanket. It establishes a multi-variable attention gate that scales weights
non-linearly using the coupled interaction term (q * v). This maps how well
queries match keys in direct proportion to how chaotic or stable the
context sequence variance is, while safely binding all final values between [-1, 1].

HOW TO USE IT (NUMPY TRANSFORMER ATTENTION STEP):

import numpy as np

def ags_power_attention_matrix(Query, Key, variance_footprint):
# 1. Compute the standard normalized Query-Key inner product matrix
d_k = Query.shape[-1]
q = np.dot(Query, Key.T) / np.sqrt(d_k)
v = variance_footprint

# 2. Execute the multi-variable cross-interaction polynomial
internal_poly = (0.312754428 * q) + (0.118784360 * (q * v)) + (0.091322486 * (q**2)) + (0.040562366 * v)

# 3. Generate the self-limiting attention mapping matrix
A = np.tanh(internal_poly)
return A
Enter fullscreen mode Exit fullscreen mode

Top comments (0)